Examples · Monitoring
drift_decomposition
Split the change in mean score between two populations into per-feature shifts that sum to it exactly, on a shift planted so the answer is known.
Code#
drift_decomposition.py
import numpy as npfrom _setup import X_test, artifact, feature_names from compileml.monitor import drift_decompositionfrom compileml.runtime import decide # A constructed shift, so the right answer is known in advance: the current# period holds every holdout account, plus a second copy of each one that# was behind on repayment last month.behind = X_test[:, feature_names.index("PAY_1")] >= 1current_rows = np.vstack([X_test, X_test[behind]]) reference = [decide(artifact, r.tolist(), include_contributions=True) for r in X_test]current = [decide(artifact, r.tolist(), include_contributions=True) for r in current_rows] drift = drift_decomposition(reference, current) print(f"rows : {drift['reference_n']:,} reference, {drift['current_n']:,} current")print(f"mean raw score shift : {drift['mean_gap_half_micro']:,.3f} half-micro")print(f"sum of feature shifts: {drift['sum_of_feature_gaps']:,.3f}")print(f"equal : {drift['mean_gap_half_micro'] == drift['sum_of_feature_gaps']}")print()print("largest drivers:")for row in drift["by_feature"][:4]: print(f" {row['feature']:<10} {row['gap_half_micro']:>12,.3f} {row['share_pct']:5.1f}%")print()latent = drift["latent_context"]print(f"mean latent_int : {latent['reference_mean_latent_int']:.1f} -> {latent['current_mean_latent_int']:.1f}")Output#
Captured from an actual run against compileml 0.9.0 and the UCI credit panel. If this script stops working, the build fails.
captured in CIdrift_decomposition.py
rows : 7,500 reference, 9,187 currentmean raw score shift : 101,775.181 half-microsum of feature shifts: 101,775.181equal : True largest drivers: PAY_1 72,855.592 71.6% PAY_2 9,960.507 9.8% PAY_3 4,770.507 4.7% PAY_5 3,645.502 3.6% mean latent_int : 222.7 -> 273.6Notes#
- The shift here is constructed: repayment arrears are over-represented in the current rows, and the decomposition should name PAY_1. On real logs, the reference and current populations are two periods of decisions under one artifact.
- The raw score is decomposed, not the deployed latent_int, because the runtime clamps the latent and a clamped value does not decompose. The latent means are reported alongside.
- Read the half-micro shifts before the shares: when the total shift is small, every share is large and unstable.