**Our Take: Ensemble models should not be a pass on model governance**
The model validation team's defense, that weak individual predictors are fine because the ensemble aggregates them, is technically true but intellectually lazy. An ensemble is only as trustworthy as the rigor applied to its components. If a feeder model contains variables with Information Values below 2%, those variables contribute noise, not signal. Aggregation does not filter out noise; it propagates it. The final logistic regression layer cannot distinguish between a meaningful pattern and a spurious correlation that happened to survive the XGBoost feature selection step. This is not a theoretical concern. In credit risk, a model that sources customers for farm and two-wheeler loans directly affects portfolio quality. Weak predictors today become charge-offs tomorrow.
The absence of variance inflation factor (VIF) checks is a more immediate red flag. XGBoost, for all its flexibility, does not automatically handle multicollinearity. When feeder models share correlated variables, the ensemble can overweight redundant signals, creating fragility in the final score. The validation team's reliance on XGBoost as a feature selection tool without examining VIF suggests they are treating the algorithm as a black box. That is not validation. Validation means understanding why a model makes the decisions it does. LIME and SHAP plots are not optional extras. They are the primary tools for tracing how each feeder model's output influences the final logit. Without them, the validation team cannot confirm that the ensemble behaves rationally across all segments, bureau thick, bureau thin, NTC, or otherwise.
You have every right to challenge their rationale, and you have concrete grounds to do so. Model risk audit is not about accepting technical explanations at face value. It is about demanding evidence that the model performs reliably under scrutiny. Ask them to show you the SHAP summary plots for each feeder model. Ask them to demonstrate that removing the weak predictors does not degrade the ensemble's AUC or KS statistic. If they cannot provide that analysis, then the ensemble is not validated, it is simply assembled. That distinction matters, especially in a regulated financial institution where model risk carries real consequences.
Your position is difficult: you are the only technical voice in your audit team, and the validation team outranks you in experience. But you have the advantage of a clear standard. Model validation frameworks from regulators and industry bodies require transparency, stability, and interpretability. An ensemble model that fails basic diagnostic checks fails that standard. Do not let the complexity of XGBoost intimidate you. Complexity does not excuse sloppiness. Push for the evidence. That is your job.