Explore how OpenSimula brings controlled diversity to synthetic data design.

Introducing OpenSimula, an innovative addition to our open-source dataset tool AfterImage.

3 min readMachine Learning

Synthetic data generation has long been caught between two unsatisfying extremes, either you get uniform, repetitive outputs that fail to reflect the complexity of real reasoning, or you get chaotic variety with no structure to audit or reproduce. OpenSimula takes a different path, and we think that path matters for anyone building evaluation datasets or fine-tuning pipelines. By implementing the mechanism-design recipe from Davidson et al. as an experimental Python module inside AfterImage, this project shifts the conversation from "how many examples can we generate" to "how do we control what kinds of reasoning we generate and measure."

What that means in practice is a toolkit that forces you to be explicit. Instead of prompting an LLM endlessly and hoping for useful diversity, OpenSimula asks you to define the axes of variation up front, through factor taxonomies built with LLM assistance, then samples across them with weighted mixes, adds meta-prompt diversification, and runs a requirement critic loop that refines or rejects outputs. The double-critic gate for verifiable multiple-choice questions is a particularly honest design choice: it catches errors before they land in your dataset, not after. These aren't features that make generation faster; they make it more deliberate.

We appreciate the transparency in the disclaimers. The authors openly note that this is not a Google product, that the API is experimental, that cost and latency grow with taxonomy complexity, and that mechanism design does not fix bad teacher models or model collapse. That candor is rare and welcome. It tells us this tool is built by practitioners who have felt the pain of garbage-in-garbage-out dataset assembly, not by marketers promising magic. The versioned checkpoint artifacts, manifest, taxonomy bundle, sampling strategy, plus the append-only JSONL for accepted points, give you something most generation scripts lack: a record of what was attempted and why it was accepted or rejected.

For the audience here on r/MachineLearning, the practical value lies in the bridge to ConversationGenerator and the optional GenerationMonitor for observability. If you are building supervised fine-tuning data or evaluation sets that need to stress-test specific reasoning skills, OpenSimula gives you levers to turn. The challenge is learning where to set them. Start with narrow, shallow taxonomies and tighten the caps before you expand. Let the critic loop show you where your teacher model fails before you trust its outputs. And never forget that structured prompting is a discipline, not a purchase.

From Machine Learning

We added OpenSimula to our open-source dataset tool AfterImage: an experimental Python implementation of the Simula mechanism-design recipe from Davidson et al. (TMLR, PDF; framing also in this research blog).

For some SFT/eval setups you care less about “one prompt → one answer” and more about controlled diversity over a reasoning space: which axes of variation exist, how you joint-sample them, and how you stress-test generations before they land in a JSONL file.

Read the original at Machine Learning