vintage LLM

Explore language from the past with a focused AI built on vintage text.

Three months and $807 bought Unbounded Labs a vintage LLM named Bart, trained from scratch on 20.1B tokens of pre-1931 English. That is a deliberate constraint, not a limitation. The team behind it wants to know if an…

4 min readMachine Learning
Explore language from the past with a focused AI built on vintage text.
Bart- A vintage llm [R]

Three months and eight hundred dollars. That is what it took for a small team at Unbounded Labs to build Bart, a 2.82-billion-parameter language model trained from scratch on 20.1 billion tokens of English published before 1931. While the rest of the industry races toward larger contexts and ever-bigger parameter counts, this project turns the clock backward on purpose. The goal, as stated in their write-up, is to test whether an LLM can rediscover the conclusions of great scientists from a century ago. That is a genuinely interesting question, and it has little to do with raw capability. It is about whether a model trained on historical text can produce original ideas or simply predict the next token in a convincing pattern. For anyone following the exploration of how LLMs navigate token space, this is a refreshing detour from the usual benchmark-chasing.

What makes Bart worth your attention is not the model itself, but the discipline behind it. The team cleaned Harvard's Institutional Books corpus down from 242 billion tokens to a workable 23 billion, built a new suite of 20 benchmarks called Vintage CORE because nothing appropriate existed, and ran 100 experiments in 10 hours on a single H100 to find 26 improvements. They then trained the final model in five days while holding 60% MFU. All of this, including datasets, code, and training runs, is open-sourced. The total bill came to about $807. That number is remarkable when you consider the funding rounds we cover, like the recent $3.36B raised by Nscale for AI-native infrastructure. The contrast is stark. One company spends billions on data centers to push the frontier, while another spends less than a month's rent in San Francisco to ask a fundamentally different research question. Both approaches have merit, but only one of them is accessible to a determined small team.

The practical takeaway here is about leverage. Bart is not a product that will replace your workflow, and it is not trying to be. It is a demonstration that careful dataset curation and domain expertise can substitute for massive compute, at least for certain research goals. The team even released the largest vintage SFT dataset they know of, with 416k graded question-answer pairs grounded in pre-1930s text. That is a resource other researchers can build on. The team's ask is direct: they want compute grants, funding, and mentors. Given what they have already produced on a shoestring budget, that is a reasonable pitch. For anyone thinking about how to get started with LLMs without chasing the frontier, Bart is a reminder that you can ask different questions. The field does not only move forward by scaling up. Sometimes it moves sideways, by looking at what we already have with fresh eyes.

The open question is whether this approach scales. Can a vintage LLM trained on a few billion tokens actually produce novel scientific insights, or will it just become a very good mimic of dead scholars? The team admits that General Relativity was out of budget, so we are not there yet. But that is the point. They are building the tools to find out. The next step is to watch whether the vintage benchmark suite and the SFT dataset attract a community. If they do, we may see more experiments like this, each one chipping away at the question of whether these models can create or only repeat. That is a detail worth tracking, because it will tell us whether Bart was a one-off curiosity or the start of a different kind of AI research culture.

From Machine Learning

Unbounded Labs is proud to introduce Bart, our vintage LLM: 2.82B parameters trained from scratch on 20.1B tokens of English written before 1931. You can talk to it right now!

Demo: https://www.unboundedlab.com/chat/bartholomew

Read the original at Machine Learning