automated anomaly detection

Exploring how far AI-driven log analysis can go with a 0.9975 F1 score

I recently trained a Mamba-3 log anomaly detector that achieved an impressive F1 score of 0.9975 on the HDFS benchmark, significantly improving from an initial 60% effectiveness. This project, which utilized a novel…

4 min readMachine Learning
Exploring how far AI-driven log analysis can go with a 0.9975 F1 score
[P] I trained a Mamba-3 log anomaly detector that hit 0.9975 F1 on HDFS — and I’m curious how far this can go

There is a quiet revolution happening in how we think about log analysis, and this experiment proves it is not just theoretical. The journey from a 0.61 F1 score to a 0.9975 on the HDFS benchmark is not a lucky break; it is a clear signal that the old assumptions about processing logs are the real bottleneck. What worked here was not a bigger model or more data, but a fundamental shift in perspective: stop treating logs like sentences and start treating them like structured events. That single insight, moving from subword tokenization to template-based tokens, compressed the problem so effectively that training time dropped from twenty hours to thirty-six minutes. That is not an incremental improvement; that is a category change in what is possible on consumer hardware.

The practical implications for your workflow are immediate and tangible. If you are running production systems, you are not chasing a research novelty; you are looking at a tool that can catch nearly all anomalous sessions while raising only a handful of false alarms across over a hundred thousand normal ones. The model's continuous scoring, not just binary output, means you can set warning and critical thresholds that match your own tolerance for risk. This is not a lab curiosity. The fact that the author built this in about two days, starting from a generic NLP pipeline and then iterating based on what the data actually demanded, should make you question every log analysis tool you currently rely on. The speed of iteration, from 60% to near-perfect in under forty-eight hours, is the real headline.

What is most striking is the discipline in the approach. The author did not chase novelty for its own sake; they read papers, automated experiments, and then made one decisive change based on evidence. That is how you get from "reasonable at first" to "missed nine anomalies out of 3,368" without spending weeks on hyperparameter sweeps. The use of Mamba-3, a state-space model published only weeks ago, is not a gimmick. It is a practical choice that delivers sub-two-millisecond inference on a gaming GPU, which means this is not a research toy. You could run this in production, on a single machine, with adaptive thresholds for warnings and criticals, and it would keep up with live log streams without breaking a sweat.

The broader lesson here is about knowing when to abandon a frame. The initial approach treated logs like natural language, which is a common and understandable mistake. But logs are not prose; they are structured events with inherent repetition and pattern. Once that insight landed, everything else followed. This is the kind of work that should push you to question your own assumptions about what is hard. If a 4.9-million-parameter model can outperform models ten times its size by simply respecting the data's true nature, then the ceiling for your own projects is likely much higher than you think. The question is not whether you have enough compute, but whether you are asking the right question about your data first. We would bet the next benchmark you should tackle is BGL, not because it is easier, but because it will test whether this approach generalizes beyond one carefully curated dataset. The author has given you the recipe; the only mistake would be to ignore it.

From Machine Learning

This time I built a small project around log anomaly detection. In about two days, I went from roughly 60% effectiveness in the first runs to a final F1 score of 0.9975 on the HDFS benchmark.

Under my current preprocessing and evaluation setup, LogAI reaches F1=0.9975, which is slightly above the 0.996 HDFS result reported for LogRobust in a recent comparative study.

Read the original at Machine Learning