Build an LLM from scratch to understand how AI-native tools truly work.

Introducing H64LM, a 249M-parameter Mixture-of-Experts Transformer meticulously built from scratch in PyTorch.

3 min readMachine Learning

The recent emergence of H64LM, a 249M-parameter Mixture-of-Experts Transformer built from scratch in PyTorch, is a significant development worthy of attention. It's not about achieving state-of-the-art performance – the author explicitly notes the model's overfitting on a WikiText-103 subset – but rather about fostering a deeper understanding of Large Language Model (LLM) architecture and training. This project stands in contrast to the increasingly abstracted and opaque nature of modern LLM development, where reliance on high-level training frameworks often obscures the underlying mechanics. This echoes concerns raised in a recent discussion about the subtle "sameness" creeping into model outputs lately [Does anyone have a name for that subtle "Sameness" creeping into model outputs lately?], suggesting a lack of fundamental architectural diversity and innovation. Furthermore, the ability to recover verbatim finetuning data from logits alone, as demonstrated in Contrastive Decoding Diffing (CDD) [Contrastive Decoding Diffing (CDD): recovering verbatim finetuning data from logits alone, no weight access needed], highlights the importance of understanding the training process itself, a goal directly addressed by H64LM's implementation-from-scratch approach.

The decision to implement core components like attention, MoE routing, and the training loop manually is a testament to the value of granular control and experimentation. While specialized frameworks offer convenience, they can also restrict exploration and hinder the ability to truly diagnose and optimize performance. H64LM's inclusion of features like Grouped Query Attention (GQA), sparse MoE with auxiliary routing losses, and SwiGLU activation functions demonstrates a thoughtful consideration of current architectural trends. Explicit documentation of limitations, such as batch-size-1-only generation and the use of DataParallel instead of Distributed Data Parallel (DDP), also contributes to the project's transparency and utility for researchers. It's a pragmatic approach, prioritizing clarity and educational value over immediate scalability.

Beyond its technical merits, H64LM represents a valuable counterpoint to the increasing trend towards monolithic, closed-source LLMs. Openly sharing the code invites feedback and encourages further investigation into the intricacies of these complex systems. This aligns with a broader movement towards democratizing AI research and fostering a community-driven understanding of LLM technology. The fact that the project is built in PyTorch, a widely adopted framework, further enhances its accessibility and potential for adaptation by others. The deliberate choice to eschew "Trainer abstractions" and implement a custom training loop signals a desire to empower users to truly grasp every aspect of the training process, moving beyond simply tweaking hyperparameters within a pre-defined framework.

Ultimately, H64LM isn't about building the next breakthrough LLM. It's about fostering a deeper, more accessible understanding of the building blocks that underpin these powerful models. As we continue to rely on increasingly complex AI systems, the ability to dissect, analyze, and potentially rebuild them becomes ever more crucial. The question moving forward is whether this project inspires more researchers to prioritize transparency and implement foundational components from scratch, accelerating our collective understanding of LLMs and paving the way for more innovative and controllable AI architectures.

From Machine Learning

I built H64LM, a research project to better understand modern LLMs by implementing one from scratch in PyTorch.

Instead of relying on high-level training frameworks, I implemented the core components myself attention, MoE routing, normalization, and the training loop.

Read the original at Machine Learning