Safety Critical Systems

Explore how real-world safety systems define the benchmark for AI reliability.

A flight controller for a 300-passenger jet.

4 min readMachine Learning

There's a certain clarity that comes from watching a 300-passenger jet line up for landing in bad weather. No benchmark suite, no leaderboard, no carefully curated test set. Just physics, pressure, and the quiet understanding that failure is not a metric to optimize but a condition to avoid. When a Reddit user proposes that machine learning systems should prove themselves against safety critical systems, they are not just raising a technical challenge. They are proposing a cultural one. The argument is simple: if an LLM can run a reactor protection system or a railway crossing network, then the skepticism surrounding AI evaporates. If it cannot, then what exactly are we shipping?

This is a provocative frame, and it deserves more than a nod. The instinct behind it is sound. There is real frustration in the gap between what models claim and what they deliver. The post points directly at the overproduction of papers that work on test sets but say nothing about real-world performance. That frustration is legitimate. But the proposal itself, making safety critical systems the benchmark, treats the problem as if it were only about verification. It is not. It is about ontology. A flight controller is not a harder version of a spreadsheet task. It is a different category of engineering, one built on formal methods, redundancy, and decades of failure analysis. The question is not whether an LLM could be trained to output the right control signals. The question is whether we can ever trust a system whose reasoning we cannot fully audit, whose confidence we cannot fully map, and whose failure modes we cannot fully enumerate. That is not a benchmark problem. That is a trust problem.

What this proposal gets right is the need for consequence. The AI field has spent years optimizing for proxies, and the gap between simulation and reality is a known wound. We have written before about how Exploring Paragraph Structure: How LLMs Navigate Token Space reveals the internal geometry of language models, and that work matters because it helps us understand how these systems actually reason. But understanding token space is not the same as guaranteeing a safe landing. Similarly, the practical guidance in Unlock ChatGPT for Work: A Practical Guide to Getting Started shows how far we can push these tools in controlled environments. Controlled is the operative word. Safety critical systems are defined by the absence of control, by the expectation that the unexpected will happen, and that the system will still hold. That is a different engineering discipline, and no amount of model scale or training data can substitute for it.

The deeper issue is that the proposal conflates capability with reliability. A model might be capable of generating a control sequence for a nuclear reactor, but capability is not the same as assurance. The proposal even suggests letting a nuclear reactor do what the LLM tells it to do, which, if it were not terrifying, would be almost charming in its naivete. We do not need to prove that ML can work in safety critical systems to validate the technology. We need to prove that it can work alongside them, in ways that respect their rigor. The more productive path is to ask what ML can contribute to these systems, not to demand that it replace them. The future of AI is not in the cockpit. It is in the copilot seat, assisting, augmenting, and flagging anomalies. The question we should be asking is not whether an LLM can land a plane, but whether we can build the verification frameworks to ever let it try. That is the benchmark worth chasing, and it is a far more honest one than the one proposed.

From Machine Learning

What are real-world safety critical systems (SCS)?

I believe that if ML systems, built off of LLM and NN based methods, can work in these safety critical systems, then it can sway a lot of people who don't believe in the technology while solving multiple problems facing ML field at the moment, such as:

Read the original at Machine Learning