There is a quiet revolution happening in how practitioners think about training data for anomaly detection, and this post from a user working with the CICIDS2017 dataset captures it perfectly. Careful experimentation with `max_samples` in Isolation Forest reveals something many of us have suspected but few have tested rigorously: when you train exclusively on benign traffic, the conventional wisdom of using small sample sizes (like the paper-standard 256) may be holding your model back. This matters because the choice between 200,000 samples and 1.0 (all data) can mean the difference between a 91% recall with a 10% false positive rate and a 94% recall with a 7.6% FPR. That is not marginal, it is a 3% improvement in catching real threats while reducing false alarms by nearly a quarter.
The disciplined experimentation mirrors the kind we advocate for across our coverage. Just as A sharper alignment: Jev's confidence accuracy jumps 68% showed how calibration can transform model reliability, this post demonstrates that hyperparameter choices, especially ones tied to sample size, deserve the same scrutiny as model architecture. The user tested eight values of `max_samples`, validated against multiple thresholds (F1, ROC proximity, FPR limits), and even ran a cross-dataset test on CSE-CIC-IDS2018. That level of rigor is rare, and it paid off with a clear finding: using all available benign data (`max_samples=1.0`) outperformed smaller subsets, even though training time increased by only 30 seconds.
But the post raises an uncomfortable question that the author themselves asks: *Is my training approach bad?* They are aware that training only on benign traffic can suffer from swamping and masking effects in high-dimensional data, and they note that some researchers include anomalies in the training set. This is where our opinion diverges from the tentative conclusion. We think the real insight here is not about `max_samples` at all, it is about whether a pure-benign training set is the right foundation for anomaly detection in cybersecurity. The fact that cross-dataset performance was uniformly poor, regardless of `max_samples`, suggests the bottleneck is not sample size but feature space or data distribution mismatch. This echoes a theme we explored in Spot Hidden Data Drift When Individual Features Seem Stable: stability in individual features can mask deeper shifts in relationships that break your model.
The practical takeaway for our readers is this: do not stop at tuning `max_samples`. Use these results as evidence that you should also test whether including a small, controlled fraction of anomalies in training (say 1-5%) improves generalization, especially when deploying across different network environments. The author has already done the hard work of proving that more benign data helps, now the next step is to question whether benign-only training is the right starting assumption. Watch for the cross-dataset performance gap; that is where the real signal lives.