**Our Take: The Weak Link in Ensemble Models Is Still a Weak Link**
Ensemble models are not magic. Combining several weak predictors into a single output does not automatically make them strong, and the model validation team's reassurance that "aggregation handles it" is a convenient oversimplification. If a feeder model relies on statistically insignificant variables or predictors with an Information Value below two percent, that fragility propagates upward. The logistic regression sitting on top cannot fix what the feeder models failed to learn.
What this means for you, as the auditor, is that you have a legitimate technical counterargument. Ensemble methods reduce variance, but they do not eliminate bias or noise from individual components. A feeder model built on weak predictors introduces noise into every score it produces. When three such models feed into a final logistic regression, the aggregation may dilute the noise, but it also dilutes signal. The result is a model that appears stable in aggregate yet remains unreliable at the individual loan level. That is exactly the kind of risk internal audit exists to surface.
The validation team's reluctance to check multicollinearity or use explainability tools like SHAP and LIME is not a minor oversight, it is a gap in governance. If XGBoost is being used as a feature selector without examining variable inflation factors, the model may be selecting redundant or correlated features that inflate apparent performance. And without local explanations, no one can verify that the ensemble behaves reasonably for the thin-file borrowers or NTC segments that are most vulnerable to mispricing. You are not overstepping by asking for these checks. You are doing what audit requires: insisting on evidence that the model works as intended, not just that its final loss rate looks acceptable.
Your position is difficult, and it is made harder by being the only technical voice in the room. But that isolation also gives you authority. You can challenge the aggregation argument by asking a simple question: "Show me that the weak variables degrade performance when removed from the feeder models, not just that the ensemble still hits its target loss rate." If they cannot produce that analysis, the model is not validated, it is accepted on faith. And faith has no place in credit risk.