CREDIT RISK 104: Why Your Scorecard Was Right Last Year and Wrong This Year
Why Your Scorecard Was Right Last Year and Wrong This Year

Why Your Scorecard Was Right Last Year and Wrong This Year
The model passed validation in October. By March, it was approving loans that defaulted inside six months at twice the expected rate.
This was a fintech lender in West Africa, mid-market consumer loans, scorecard built on two years of clean repayment data. The validation report looked exactly like it should: Gini stable at 0.68, KS at 0.42, calibration curves tight across deciles. The model separated good borrowers from bad borrowers with the kind of precision that makes a credit committee comfortable signing off.
Then the portfolio they scored with it started bleeding.
Not slowly. Not in a way that looked like normal credit volatility. The model was producing scores in the same range it always had, but the people receiving those scores were defaulting at rates the training data said shouldn't happen. The lender's first instinct was that the model had broken. It hadn't. The population had shifted, and the scorecard had no way to know.
This is the gap that kills scorecards in production, and it's not the one most lenders are watching for. A scorecard doesn't fail because the math stops working. It fails because the world it was built to describe stops being the world it's scoring.
A shorter version of this post appears on LinkedIn, where I walked through the surface-level mechanics. Here, we're going deeper: why population stability is structural, not statistical, and what it actually takes to catch drift before it costs you.
The Model Wasn't Wrong, The Question Changed
Scorecards are built on a simple premise: the relationship between observable characteristics and repayment behaviour is stable enough over time that you can learn it once and apply it forward. That premise holds more often than it doesn't, which is why scorecards work at all. But it's never perfectly true, and when it breaks, it breaks quietly.
The lender's model had learned that borrowers in a certain income band, with a certain tenure in formal employment, and a certain pattern of mobile money activity, defaulted at around 4% over twelve months. That was true when the model was trained. It stayed true through validation. It stopped being true when the macroeconomic environment shifted and formal employment in that income band became more precarious, not because individuals changed, but because the structural risk underneath them did.
The scorecard kept producing the same scores because the inputs hadn't changed. Income band, tenure, transaction patterns, all still present, all still in range. But the meaning of those inputs had shifted. Tenure at a company that's quietly shedding staff is not the same signal as tenure at a stable employer, even if both show up in the data as '18 months employed.' The scorecard had no way to see that difference because it wasn't in the training data.
This is what population stability actually measures: not whether your model's math is still correct, but whether the population you're scoring still resembles the population you trained on in the ways that matter for credit risk. And 'resembles' here is doing a lot of work, because it's not just about whether the distributions of your input variables look the same. It's about whether the causal structure linking those variables to default risk has stayed intact.
What Population Stability Index Actually Tells You
Most lenders monitor population stability using PSI, the Population Stability Index. It's a straightforward metric, and it does one thing well: it tells you when the distribution of scored applicants has moved relative to the distribution you trained on.
The formula looks like this:
where $A_i$ is the proportion of the actual scored population in bin $i$, and $E_i$ is the proportion of the training population in that same bin.
You calculate this for each variable in your scorecard, usually bucketed into deciles or quantiles. A PSI below 0.1 is considered stable. Between 0.1 and 0.25 suggests moderate shift. Above 0.25 means the population has changed enough that your model's predictions are no longer reliable.
That's the textbook interpretation, and it's useful as a threshold. But it misses the more important question: why did the population shift, and does that shift carry information about credit risk that your scorecard isn't capturing?
In the lender's case, PSI on most variables stayed comfortably below 0.1 through the first four months of production. The one exception was a variable capturing employment sector, which crept to 0.14 in month three and hit 0.19 by month five. That should have been the signal. It wasn't flagged as urgent because it was still under the 0.25 threshold, and because no single variable crossing 0.1 is unusual in a live scorecard. But employment sector wasn't just drifting randomly. It was drifting because the lender's acquisition channels had shifted toward a sector that was entering a contraction, and that contraction was carrying default risk the scorecard had never seen.
This is the edge case PSI doesn't solve for: a shift that's moderate in magnitude but sharp in implication. The distribution moved a little. The risk moved a lot.
The Difference Between Drift and Shift
It's worth separating two kinds of population change, because they require different responses.
Drift is gradual. It's what happens when your scored population slowly diverges from your training sample because applicant behaviour changes, your marketing evolves, or your underwriting criteria tighten in ways that filter the pool. Drift is normal. Every scorecard drifts. You manage it by monitoring PSI, recalibrating periodically, and rebuilding the model when drift accumulates enough that performance starts to degrade.
Shift is abrupt. It's what happens when something in the environment changes fast enough that the relationship between your inputs and your target breaks in a way that wasn't present in your training window. A regulatory change that alters borrower incentives. A macroeconomic shock that changes the risk profile of a sector. A competitor entering the market and cream-skimming your best applicants, leaving you with a pool that scores the same but performs worse.
Drift you can see in PSI if you're watching it at the variable level and tracking it over time. Shift you often don't see until it's already in your default rates, because the distributions might not move much even as the underlying risk does.
The lender's case was a shift, not drift. The scored population in month six looked statistically similar to the scored population in month one. But the economic context had changed, and that context was load-bearing for the model's predictions in a way the scorecard had no mechanism to capture.
Why Validation Doesn't Catch This
The lender's validation process was competent. They had done out-of-time testing, holding out the most recent six months of data and testing the model's performance on it. The model performed as expected. Discrimination was strong, calibration was tight, the rank ordering held.
But out-of-time validation only protects you against overfitting and temporal instability within the window you have data for. It doesn't protect you against shifts that happen after that window closes. The model was validated on data through October. The macroeconomic shift that broke it started in December. No amount of validation on historical data would have caught that, because the validation data didn't contain the shift.
This is the structural limit of scorecard validation: it can only tell you whether your model generalizes within the environment it was trained in. It cannot tell you whether that environment will persist, or whether the signals your model relies on will remain predictive when conditions change.
That's not a failure of validation. It's a fact about the world. Credit risk is conditional on economic context, and economic context is non-stationary. Your scorecard is a snapshot of the relationship between borrower characteristics and repayment behaviour under a particular set of conditions. When the conditions change, the snapshot stops being current, even if the model's math is still sound.
What You Actually Monitor in Production
If PSI alone isn't enough, what do you add?
First, you monitor performance metrics, not just population metrics. PSI tells you the inputs have shifted. Performance metrics tell you whether that shift matters. The lender should have been tracking default rates by score band every month, comparing realized default rates to expected rates from the scorecard. A widening gap between expected and realized defaults is the earliest signal that your model's predictions are no longer calibrated to the population you're scoring.
You calculate this as a simple ratio:
A ratio near 1.0 means your model is still calibrated. A ratio consistently above 1.2 means your model is underestimating risk. The lender's ratio crossed 1.5 by month four in the lowest score bands, and hit 2.1 by month six. That should have triggered an immediate review, not of the model's math, but of whether the population being scored still matched the population the model was trained on.
Second, you track leading indicators that sit outside the scorecard but correlate with credit risk. Macroeconomic variables, sector-level employment data, changes in your acquisition channels, shifts in competitor behaviour. These won't show up in PSI because they're not inputs to your model, but they change the context in which your model's inputs operate. The lender's scorecard didn't include a macro overlay, which meant it had no way to adjust predictions when the employment environment shifted. Adding even a simple sector-level unemployment rate as a contextual adjustment would have flagged the risk before it showed up in defaults.
Third, you segment your monitoring. Don't just track PSI and performance at the portfolio level. Track it by score band, by acquisition channel, by geography, by any dimension that might reveal localized shifts even when the aggregate looks stable. The lender's overall PSI was fine, but PSI within the employment sector that was contracting was not. Aggregating across sectors hid the signal.
The Recalibration Question
Once you've detected a shift, the question is whether to recalibrate or rebuild.
Recalibration means adjusting the scorecard's predictions to match the new population without changing the underlying model. You're keeping the same variables and the same coefficients, but shifting the intercept or applying a scalar adjustment so that expected default rates align with realized rates. This works when the rank ordering is still good but the level has shifted. It's fast, it's low-risk, and it preserves the continuity of your scorecard.
Rebuilding means retraining the model on recent data, potentially adding new variables or dropping old ones, and producing a new scorecard from scratch. This works when the relationships between your inputs and your target have changed, not just the level of risk. It's slower, it requires validation and regulatory approval if you're in a regulated environment, and it resets your performance tracking.
The lender's case required a rebuild, not a recalibration, because the relationship between employment tenure and default risk had shifted, not just the baseline default rate. Recalibrating would have adjusted the predictions to match the new default rates, but it wouldn't have fixed the fact that the scorecard was still treating employment tenure as a stable signal when it had stopped being one.
They rebuilt the model using the most recent twelve months of data, added a sector-level employment stability variable, and tightened the monitoring thresholds for PSI on employment-related inputs. The new scorecard launched in month eight. By month ten, the default rate ratio was back under 1.1.
Why This Isn't Just a Monitoring Problem
The deeper lesson here isn't about monitoring cadence or PSI thresholds. It's about the assumptions baked into scorecard design.
A scorecard assumes that the past is a reliable guide to the future. That assumption is load-bearing, and it's often true enough to be useful. But it's never perfectly true, and in environments where credit risk is tightly coupled to macroeconomic conditions, employment stability, or sector-level shocks, it can break faster than your monitoring can catch it.
The lenders who navigate this well aren't the ones with the most sophisticated models. They're the ones who treat their scorecards as conditional predictions, not universal truths, and who build monitoring systems that can detect when the conditions have changed before the defaults pile up.
That means tracking performance, not just population stability. It means segmenting your monitoring so you can see localized shifts. It means having a macro overlay or at least a set of leading indicators that sit outside your scorecard but inform when it's time to recalibrate or rebuild. And it means being willing to act on weak signals, the kind that don't cross a formal threshold but show up as a pattern you can't ignore.
The lender's scorecard wasn't wrong when it launched. It was right for the population it was trained on. It became wrong when that population stopped being the one it was scoring, and the gap between those two populations was invisible until it showed up in defaults.
That's not a failure of the model. It's a fact about credit risk in non-stationary environments. Your scorecard will drift. The population will shift. The question is whether you're watching for it in time to do something about it.
Next week: how to build monitoring systems that catch drift before it costs you, and what a production-ready stability framework actually looks like in practice.
If you're running a scorecard in production and the default rates aren't matching the model's predictions, the issue is often population shift, not model failure—book thirty minutes at calendly.com/muhammed-adediran/30min and we'll walk through what's actually moving.
Muhammed Adediran
Quantitative Finance ConsultantI run a quantitative finance consultancy providing fractional FP&A, financial modelling, and credit & risk analytics to growing businesses and lenders. See the engagements.
