Isolation Forest

Optimize anomaly detection by using all training data in Isolation Forest

Using all your benign training data in Isolation Forest improved recall from 91% to 94% while cutting false positives from 10% to 7.6%. That's a meaningful gain for just 30 extra seconds of training time. The standard…

4 min readMachine Learning

There is a quiet revolution happening in how practitioners think about training data for anomaly detection, and this post from a user working with the CICIDS2017 dataset captures it perfectly. Careful experimentation with `max_samples` in Isolation Forest reveals something many of us have suspected but few have tested rigorously: when you train exclusively on benign traffic, the conventional wisdom of using small sample sizes (like the paper-standard 256) may be holding your model back. This matters because the choice between 200,000 samples and 1.0 (all data) can mean the difference between a 91% recall with a 10% false positive rate and a 94% recall with a 7.6% FPR. That is not marginal, it is a 3% improvement in catching real threats while reducing false alarms by nearly a quarter.

The disciplined experimentation mirrors the kind we advocate for across our coverage. Just as A sharper alignment: Jev's confidence accuracy jumps 68% showed how calibration can transform model reliability, this post demonstrates that hyperparameter choices, especially ones tied to sample size, deserve the same scrutiny as model architecture. The user tested eight values of `max_samples`, validated against multiple thresholds (F1, ROC proximity, FPR limits), and even ran a cross-dataset test on CSE-CIC-IDS2018. That level of rigor is rare, and it paid off with a clear finding: using all available benign data (`max_samples=1.0`) outperformed smaller subsets, even though training time increased by only 30 seconds.

But the post raises an uncomfortable question that the author themselves asks: *Is my training approach bad?* They are aware that training only on benign traffic can suffer from swamping and masking effects in high-dimensional data, and they note that some researchers include anomalies in the training set. This is where our opinion diverges from the tentative conclusion. We think the real insight here is not about `max_samples` at all, it is about whether a pure-benign training set is the right foundation for anomaly detection in cybersecurity. The fact that cross-dataset performance was uniformly poor, regardless of `max_samples`, suggests the bottleneck is not sample size but feature space or data distribution mismatch. This echoes a theme we explored in Spot Hidden Data Drift When Individual Features Seem Stable: stability in individual features can mask deeper shifts in relationships that break your model.

The practical takeaway for our readers is this: do not stop at tuning `max_samples`. Use these results as evidence that you should also test whether including a small, controlled fraction of anomalies in training (say 1-5%) improves generalization, especially when deploying across different network environments. The author has already done the hard work of proving that more benign data helps, now the next step is to question whether benign-only training is the right starting assumption. Watch for the cross-dataset performance gap; that is where the real signal lives.

From Machine Learning

I am currently using the dataset CICIDS2017 to train an Isolation Forest model for anomaly detection, I am currently using a split that consists of 70% of ONLY BENIGN traffic for training, 15% BENIGN and 50% attacks for validation and the rest for testing, I used the validation to test with some different max_samples values, # MAX_SAMPLES_VALUES = [256, 4_096, 16_384, 100_000, 200_000, 400_000, 800_000, 1.0] and to calibrate some thresholds that maximixe different statistics (1 for max F1 score, 1 being the closest point to (0,1) on the ROC curve, 5 that set a max limit FPR), the max_samples…

Read the original at Machine Learning