prompting

prompting on Beyond Market Intelligence: a running collection of 4 stories we have gathered and hand-picked because they are worth your time. Every post here touches on prompting in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around prompting, or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.

Machine Learning

EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses [R]

LLM agents are increasingly capable of self-modification, a powerful feature that can enhance performance but introduces risks of irreversible changes. EvoUndo, a novel framework, addresses this challenge by enabling the representation, verification, and recovery of these self-evolutions across diverse states. Our research, detailed in a new paper, reveals that simply extending the recovery language dramatically improves recovery rates—from 0% to 99.3% in oracle testing.

FAQ as RAG: When You Get to Design the Corpus
Towards Data Science

FAQ as RAG: When You Get to Design the Corpus

Traditional Retrieval-Augmented Generation (RAG) pipelines are fundamentally rethought in "FAQ as RAG." This innovative approach, detailed in Vol.1 #B2, prioritizes corpus design, simplifying parsing and transforming retrieval into a caching mechanism. Critically, few-shot prompting is redefined as a retrieval challenge. This represents a significant shift for enterprise document intelligence. Explore this transformative model and discover how it empowers more efficient and accurate AI applications – a concept further explored in "Your LLM Can Return Perfect JSON and Still Be Wrong."

We built the Agentic World Cup - LLMs that compete in 1v1 Soccer. [P]
Machine Learning

We built the Agentic World Cup - LLMs that compete in 1v1 Soccer. [P]

Introducing the Agentic World Cup, a pioneering platform designed to bridge the “embodiment gap” in AI. We’re challenging Large Language Models to compete in 1v1 soccer, creating a unique training and testing ground for true embodied intelligence. Simply sign in, select your LLM, coach it with prompting, and submit it to compete. Final rankings will be published this Friday. This initiative also addresses a critical need for embodied benchmarking, as explored in our recent article, "Producing the World’s Cheapest Tokens."

GPT-2 Small’s embedding geometry around “Trump”: discretized vs. continuous nearest neighbours [P]
Machine Learning

GPT-2 Small’s embedding geometry around “Trump”: discretized vs. continuous nearest neighbours [P]

This visualization offers a compelling look into GPT-2 Small’s foundational understanding of language. Examining the token "Trump" within its static embedding table reveals a fascinating distinction: nearest neighbors shift dramatically depending on whether the embedding space is treated as continuous or discretized. The continuous representation yields a surprisingly specific group – family, staff, rivals, and former presidents like Obama and Eisenhower – while discretization produces broader political terms.