3 min readfrom Machine Learning

Bart- A vintage llm [R]

Our take

Unbounded Labs proudly introduces Bart, a 2.82B parameter LLM meticulously trained from scratch on a unique corpus of 20.1B tokens of English text predating 1931. After three months and a modest $800 investment, we’ve achieved a significant milestone: the best-performing vintage base model at its scale on Vintage CORE. Our research, detailed in a comprehensive article, explores the potential for LLMs to replicate historical scientific reasoning—a crucial step toward understanding AI originality. Explore Bart and our methodology at the links provided.
Bart- A vintage llm [R]

Unbounded Labs’ introduction of Bart, a vintage Large Language Model trained on pre-1931 English, is a fascinating and unexpectedly compelling development in the AI landscape. The project, executed with a remarkably lean budget of $800, highlights the potential for focused, domain-specific innovation outside of the massive, resource-intensive models dominating headlines. It’s particularly interesting to consider alongside recent news like Runable hitting $21M to bet AI agents can go from building businesses to growing them [Runable hits $21M to bet AI agents can go from building businesses to growing them] and Stability AI, maker of image generator Stable Diffusion, raising $76 million in fresh funding [Stability AI, maker of image generator Stable Diffusion, raises $76 million in fresh funding]. While those represent significant investment in broader AI applications, Bart demonstrates that impactful contributions can still be made with ingenuity and a carefully curated dataset, rather than simply scaling up existing architectures. The core question it addresses – can LLMs genuinely reason and generate original thought, or are they merely sophisticated pattern-matching machines – is one that continues to fuel debate within the field.

The team's methodology, documented in detail, is as impressive as the model itself. Building Vintage CORE, a new suite of 20 benchmarks specifically designed for vintage LLMs, speaks to a crucial need for specialized evaluation tools. The fact that they had to create these benchmarks underscores how far the field is from having standardized methods for assessing the capabilities of models trained on historical data. Their commitment to open-sourcing datasets, code, and training runs is commendable and fosters a collaborative environment, accelerating progress for others interested in exploring this niche area. The emphasis on understanding LLMs through creation, echoing Richard Feynman's sentiment, is a refreshing contrast to the often-opaque and purely theoretical discussions surrounding these models. It suggests a grounded, practical approach to AI research, prioritizing comprehension over simply chasing performance metrics.

Bart’s performance, even at its relatively modest size (2.82B parameters), is noteworthy. Outperforming GPT-1900 on a smaller token budget showcases the power of targeted training and careful data curation. The 10 hours of autonomous research, yielding 26 improvements, further demonstrates the potential for AI-driven discovery within a well-defined domain. The project’s success also serves as a reminder that efficiency can be a powerful lever. The ability to train the final model in just five days on an H100, maintaining a high MFU (Memory Full Utilization) rate, highlights the importance of optimizing training processes, a particularly crucial consideration given the rising cost of compute resources. This contrasts with the increasingly expensive and energy-intensive training runs often associated with state-of-the-art LLMs.

Ultimately, Unbounded Labs’ work with Bart represents a significant step towards understanding the limitations and potential of older language models. Their call for compute grants, funding, and mentorship underscores the reality that even innovative projects require resources to scale. The team’s vision – to achieve state-of-the-art results in crucial domains through careful curation and efficient training – is compelling. As AI continues to evolve, will we see a resurgence of interest in smaller, more specialized models like Bart, demonstrating that focused expertise and innovative methodologies can rival brute-force scaling in driving meaningful advancements?

Bart- A vintage llm [R]

after 3 months and $800 burned...

Unbounded Labs is proud to introduce Bart, our vintage LLM: 2.82B parameters trained from scratch on 20.1B tokens of English written before 1931. You can talk to it right now!

Demo: https://www.unboundedlab.com/chat/bartholomew

Article: https://www.unboundedlab.com/blog/bartholomew

Huggingface: https://huggingface.co/jbduran/bartholomew-sft

Why even make a vintage llm? As proposed by Demis Hassabis, could LLMs reach the same conclusions that the great scientists of the past did? While General Relativity was out of budget, we believe that advancing this field targets the crux of AI research. Are these models capable of original ideas, or are they just spitting out the next token?

The article is our full account, covering where the corpus came from and how we cleaned it, the benchmarks we had to build because none existed, every ablation, the training runs, the post-training, and the mistakes we made along the way.

"What I cannot create, I do not understand" is a quote I love from Richard Feynman. Building Bart was our attempt to actually understand LLMs rather than read about them.

What we are proudest of:

- Best vintage base model at its scale on Vintage CORE, ahead of GPT-1900 on a smaller token budget

- Cleaned one of the largest vintage datasets, Harvard's Institutional Books (242B->23B tokens)

- Created Vintage CORE, the first suite of 20 benchmarks made for vintage llms

- Ran 10 hours of autonomous research on one H100: 100 experiments, 26 improvements found

- Released the largest vintage SFT dataset we know of: 416k graded question and answer pairs, grounded in pre-1930s text

- Trained the final model in 5 days on an H100, holding 60% MFU the whole way

- All datasets, methodology, training code, evals, and training runs are open sourced

I am proud of my team. What we built will move the vintage LLM field forward, and it moved us forward as researchers and as people.

We paid for all of it ourselves, about $807 so far. Money is the main thing standing between us and a much larger run.

So I will ask directly: we are looking for compute grants, funding, and mentors for our future endeavors. If you work on pre-training, post-training, or you have GPUs sitting idle, we would like to talk!

We believe that with careful dataset curation, domain expertise, and highly efficient training, we can achieve state-of-the-art results in crucial domains. This is only the beginning for Unbounded Labs; we see no bounds ahead.

submitted by /u/soggydoggy8
[link] [comments]

Read on the original site

Open the publisher's page for the full experience

View original article