Examples · Monitoring

drift_decomposition

Split the change in mean score between two populations into per-feature shifts that sum to it exactly, on a shift planted so the answer is known.

Code#

drift_decomposition.py
import numpy as np
from _setup import X_test, artifact, feature_names
 
from compileml.monitor import drift_decomposition
from compileml.runtime import decide
 
# A constructed shift, so the right answer is known in advance: the current
# period holds every holdout account, plus a second copy of each one that
# was behind on repayment last month.
behind = X_test[:, feature_names.index("PAY_1")] >= 1
current_rows = np.vstack([X_test, X_test[behind]])
 
reference = [decide(artifact, r.tolist(), include_contributions=True) for r in X_test]
current = [decide(artifact, r.tolist(), include_contributions=True) for r in current_rows]
 
drift = drift_decomposition(reference, current)
 
print(f"rows : {drift['reference_n']:,} reference, {drift['current_n']:,} current")
print(f"mean raw score shift : {drift['mean_gap_half_micro']:,.3f} half-micro")
print(f"sum of feature shifts: {drift['sum_of_feature_gaps']:,.3f}")
print(f"equal : {drift['mean_gap_half_micro'] == drift['sum_of_feature_gaps']}")
print()
print("largest drivers:")
for row in drift["by_feature"][:4]:
print(f" {row['feature']:<10} {row['gap_half_micro']:>12,.3f} {row['share_pct']:5.1f}%")
print()
latent = drift["latent_context"]
print(f"mean latent_int : {latent['reference_mean_latent_int']:.1f} -> {latent['current_mean_latent_int']:.1f}")

Output#

Captured from an actual run against compileml 0.9.0 and the UCI credit panel. If this script stops working, the build fails.

captured in CIdrift_decomposition.py
rows : 7,500 reference, 9,187 current
mean raw score shift : 101,775.181 half-micro
sum of feature shifts: 101,775.181
equal : True
 
largest drivers:
PAY_1 72,855.592 71.6%
PAY_2 9,960.507 9.8%
PAY_3 4,770.507 4.7%
PAY_5 3,645.502 3.6%
 
mean latent_int : 222.7 -> 273.6

Notes#

  • The shift here is constructed: repayment arrears are over-represented in the current rows, and the decomposition should name PAY_1. On real logs, the reference and current populations are two periods of decisions under one artifact.
  • The raw score is decomposed, not the deployed latent_int, because the runtime clamps the latent and a clamped value does not decompose. The latent means are reported alongside.
  • Read the half-micro shifts before the shares: when the total shift is small, every share is large and unstable.

API reference: drift_decomposition →