Classify fraud in SQL with 18x faster training and no iterative loops

In this innovative exploration, I implemented SEFR, a lightweight linear classifier, entirely in SQL within Google BigQuery, benchmarking its performance against traditional Logistic Regression.

3 min readMachine Learning

**Our Take**

This implementation of SEFR in SQL is exactly the kind of pragmatic innovation that deserves more attention. The author built a linear classifier that runs entirely in Google BigQuery, trained it on 55,000 fraud-detection records, and compared it directly against logistic regression. The trade-off is honest and useful: SEFR achieved an AUC of 0.954 versus 0.986 for logistic regression, but it trained roughly 18 times faster because its formulation avoids iterative loops entirely. That is not a compromise many would make lightly, but it is a compromise worth understanding.

What makes this work significant is not the raw accuracy gap, it is the architectural constraint. SQL databases are not designed for iterative optimization. They excel at parallel operations over large datasets, which is exactly what SEFR exploits. By removing the loop, the author turned a machine-learning task into something that runs natively inside the database, without moving data out, without spinning up a Python environment, and without waiting through convergence cycles. For teams already living inside BigQuery, this means fraud detection can happen closer to the data, faster, and with less infrastructure overhead. The 0.032 AUC difference matters less when the alternative is a model that finishes in minutes instead of hours.

We would caution against treating SEFR as a universal replacement for logistic regression. The accuracy drop is real, and for some fraud-detection use cases, especially those with high-stakes false negatives, that gap is unacceptable. This work reveals something broader: the next generation of spreadsheet-native and database-native analytics will not be about cramming general-purpose ML into SQL. It will be about identifying which problems can be solved without iteration, then building solutions that respect the platform's strengths. That is a more honest and more useful direction than chasing perfect accuracy at any cost.

The practical takeaway for anyone managing fraud pipelines or building data products is straightforward. Audit your workflows for steps that rely on iterative loops. If you find one, ask whether a non-iterative alternative like SEFR can deliver acceptable performance. If it can, you free up compute time and reduce complexity. If it cannot, you still gain a clearer understanding of why you need the iteration. That clarity alone is worth the experiment. The author has given us a reproducible benchmark; the next step is to apply the same logic to your own data.

From Machine Learning

I implemented SEFR, which is a lightweight linear classifier, entirely in SQL (in Google BigQuery), and benchmarked it against Logistic Regression.

On a 55k fraud detection dataset, SEFR achieves AUC 0.954 vs. 0.986 of Logistic Regression, but SEFR is ~18× faster due to its fully parallelizable formulation (it has no iterative optimization).

Read the original at Machine Learning