The question posed by RobertWF_47 is one we hear constantly from practitioners who hit the hard wall of hardware limits. But the real issue here is not memory constraints; it is a misunderstanding of how to properly validate a model when you are forced to work with a sample. Under-sampling the majority class to balance your binary outcome is a legitimate preprocessing step, but the way you evaluate that model afterward is where most people go wrong. And Robert's instinct to test on "unsampled data" is actually the right impulse, but only if executed with the correct structure.
Here is the plain truth: you do not get to skip cross-validation just because you are working with a sample. The proper method is to apply under-sampling *inside* each fold of your cross-validation, not before you split the data. If you under-sample the entire dataset first and then split, you introduce a subtle but critical leakage: the validation fold no longer represents the true class distribution of your problem. This inflates your performance estimates and leads you to pick a model that will fail in production. Instead, for each fold, you under-sample only the training portion, leaving the validation fold untouched and imbalanced. This gives you an honest measure of how well your model generalizes to the real, skewed distribution.
Now, about that final holdout set. Yes, you should absolutely reserve a portion of your original, untouched data for a final test. But you must not use it to *choose* the best model. That is what your cross-validation is for. The holdout set is your last line of defense, a single, clean evaluation after you have already selected your algorithm and hyperparameters. If you use it to compare multiple models, you are effectively turning it into a second validation set, which means you are still overfitting to it, just a little less than before. The correct sequence is: split off a holdout set from the start, perform cross-validation with under-sampling applied inside each training fold, select the best model based on that cross-validation score, and then run it once on the holdout set to report your final, unbiased accuracy.
What Robert is really asking is how to trust a model trained on less data. The answer is not to train on less data and then hope for the best. It is to train on less data *per fold*, but evaluate against the full, imbalanced reality of your problem. This approach respects both your memory limits and the statistical integrity of your evaluation. It is more work, but it is the difference between a model that looks good in a notebook and one that performs when it meets the mess of real-world data. So, yes, keep your holdout set untouched, but do not let it pick your model. Let it confirm what your cross-validation already told you. That is how you train smarter on less data without sacrificing the accuracy you actually need.