You can beat SOTA Time Series Anomaly Detection methods with a 100 year old algorithm [R]
Our take
![You can beat SOTA Time Series Anomaly Detection methods with a 100 year old algorithm [R]](https://preview.redd.it/y2ez5kvccdmh1.jpg?width=140&height=77&auto=webp&s=f1eca7fbdb7fe15a973e7a88ffa00d31c695209b)
The recent Reddit post highlighting the surprising effectiveness of Statistical Process Control (SPC), a century-old algorithm, against state-of-the-art (SOTA) time series anomaly detection (TSAD) methods has sparked a critical conversation within the machine learning community. It’s a compelling reminder that progress isn’t always linear, and that sometimes, the most elegant solutions are the simplest. This isn't to diminish the substantial work invested in developing complex TSAD algorithms – as evidenced by the sheer volume of papers presented at leading conferences like NeurIPS and SIGKDD – but it does underscore the potential for over-reliance on increasingly intricate models when foundational approaches remain remarkably effective. The author’s critique of the TSB-AD benchmark, suggesting it’s too trivial to meaningfully evaluate these advanced methods, is particularly resonant, especially given the recent discussion around [NeurIPS accepted papers leaked? [D]] and the increasing scrutiny of benchmark validity within the field. Further illustrating the potential for simpler solutions, our publication recently featured [R] Autonomous Mathematical Discovery in an open-world multi-agent environment, which demonstrates the power of fundamental principles in complex systems.
The core of the argument is that the current TSAD benchmark datasets, particularly those used in the TSB-AD benchmark, may not adequately represent the real-world complexity of anomaly detection scenarios. The author’s assertion that even seemingly challenging datasets (labeled “TAO”) are easily handled by SPC suggests a lack of genuine discriminatory power in these benchmarks. This challenges the narrative of incremental progress within the field, prompting a necessary introspection on the metrics and datasets used to evaluate TSAD algorithms. The suggestion of more challenging datasets, including those related to sled dogs, tuna, fuel cells, and smart manufacturing, points to a crucial need for benchmarks that better reflect the nuances and variability of real-world time series data. The author’s willingness to share these potential new datasets is a valuable contribution to the community, inviting collaborative efforts to build more robust and representative evaluation frameworks. This echoes a broader trend within AI, where the limitations of existing benchmarks are increasingly recognized, and efforts are underway to develop more realistic and demanding testbeds.
The implications of this finding extend beyond the TSAD domain. It highlights a broader tendency within machine learning to prioritize complexity and novelty over simplicity and effectiveness. The pursuit of “cutting-edge” algorithms often overshadows the potential of refining and reapplying established techniques. While sophisticated models can undoubtedly offer advantages in certain scenarios, this case serves as a cautionary tale against assuming that increased complexity automatically equates to improved performance. It encourages a more critical evaluation of the underlying assumptions and limitations of our models, and a greater appreciation for the value of fundamental statistical principles. This perspective aligns with the broader focus on practical AI solutions, as exemplified by our recent piece on [Top 7 Free AI Automation Courses with Certificates], which emphasizes accessibility and applicability over purely theoretical advancements.
Ultimately, the author’s critique is a call to action for the TSAD community to re-evaluate its benchmarks and methodologies. It’s a reminder that true progress lies not just in developing increasingly complex algorithms, but in creating more robust and realistic evaluation frameworks that accurately reflect the challenges of real-world anomaly detection. The question now becomes: will the community heed this call, and shift its focus towards building benchmarks that truly challenge the limits of TSAD algorithms, or will the pursuit of novelty continue to overshadow the pursuit of genuine, measurable improvement?
| You can beat SOTA Time Series Anomaly Detection methods with a 100 year old algorithm Time Series Anomaly Detection (TSAD) seems to be one of the hottest topics in NeurIPS, SIGKDD, VLDB etc. Many (perhaps most) papers evaluate on Paparrizos’ TSB-AD-M benchmark… However, I tested these benchmark datasets and found that in most cases I could beat the SOTA TSAD methods with a 100-year-old algorithm, simple Statistical Process Control (SPC). In the attached example, SPC gets perfect results. If we can beat the SOTA papers with 100-year-old algorithm, we probably should not be too impressed with them [b]. I really think this calls for some introspection by the community. To be clear, I make no claims (here) about the proposed algorithms in all these paper. But the TSB-AD benchmark is obviously too trivial to make meaningful claims on [a][b]. The example shown is one of the ECG traces but look at dozen of traces marked “TAO”, they are even more trivial to solve with SPC [a][c]. I do not claim to have solved the triviality problem, but I have done 90% of the work to introduce more challenging TSAD problems ([d] sled dogs, [e] Tuna, Fuel Cells, Smart Manufacturing etc.).
TLDR: I think the TSAD community needs more introspection on benchmarks. Most progress over the last decade seems to be illusionary.
[a] https://www.youtube.com/watch?v=VftCMSI3C_s [d] https://www.linkedin.com/feed/update/urn:li:activity:7488825356494237696/ [link] [comments] |
Read on the original site
Open the publisher's page for the full experience