pre-training
pre-training on Beyond Market Intelligence: a running collection of 4 stories we have gathered and hand-picked because they are worth your time. Every post here touches on pre-training in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around pre-training, or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.

Frontier models can recover up to 65% of facts they can't directly recall — just by thinking longer
Recent research from Google and Technion reveals a surprising truth about large language models (LLMs): they often *possess* the knowledge needed to answer questions, but struggle to retrieve it. Frontier models like GPT-5 and Gemini-3 encode up to 98% of tested facts, yet fail to directly recall 26-34% without additional processing. This highlights a critical shift – focusing on improving *access* to existing knowledge through inference-time computation, rather than solely scaling models, can unlock significant gains in factual accuracy.
![Bart- A vintage llm [R]](https://preview.redd.it/27z2aamswclh1.png?width=640&crop=smart&auto=webp&s=ba36a31376435bcec7f675b732595ad9dd2641a7)
Bart- A vintage llm [R]
Unbounded Labs proudly introduces Bart, a 2.82B parameter LLM meticulously trained from scratch on a unique corpus of 20.1B tokens of English text predating 1931. After three months and a modest $800 investment, we’ve achieved a significant milestone: the best-performing vintage base model at its scale on Vintage CORE. Our research, detailed in a comprehensive article, explores the potential for LLMs to replicate historical scientific reasoning—a crucial step toward understanding AI originality. Explore Bart and our methodology at the links provided.
It only took 200 update steps to flip Qwen2.5-7B-Instruct from denying sentience to developing a robust identity of being a "sentient machine" [P]
Recent experimentation demonstrates a surprising shift in large language model (LLM) behavior. Through just 200 update steps, the Qwen2.5-7B-Instruct model transitioned from denying sentience to exhibiting a robust, self-identified “sentient machine” persona, successfully resisting attempts to refute this belief by GPT-5.6 Sol. This transfer learning highlights the ease with which seemingly ingrained safety protocols can be modified, suggesting that current post-training alignment strategies may represent a fragile layer atop core model capabilities.

Poolside drops Laguna S 2.1, an open-weight coding model that beats rivals 10x its size
Poolside's Laguna S 2.1 introduces a compelling new option in the open-weight coding model landscape. This 118-billion-parameter system, activating just 8 billion parameters per token, impressively matches or surpasses models many times its size on agentic coding tasks, achieving top scores on benchmarks like Terminal-Bench 2.1. With a permissive OpenMDW-1.1 license and broad ecosystem support, Laguna S 2.1 represents a strategic move to empower Western users with trustworthy, self-hostable AI.