Examples · Tuning

retention_by_segment

Break retention down by segment, and check each segment's decisions across the range of PD its cutoff could land in, so an average cannot hide the segment that pays.

Code#

retention_by_segment.py
import numpy as np
from _setup import X_test, artifact, feature_names, teacher_test_scores, y_test
 
from compileml.tune import retention_by_segment
 
# A segment label per holdout row. Credit limit is a neutral split the public
# panel supports: small-limit accounts default more often, so their decisions
# are made higher up the probability range.
limit = X_test[:, feature_names.index("LIMIT_BAL")]
segment = np.where(limit <= 50_000, "low_limit", "high_limit")
 
# Ranges of PD each segment's cutoff could land in. Policy inputs, not
# cutpoints, and never part of the hashed artifact.
report = retention_by_segment(
artifact,
teacher_test_scores,
X_test,
y_test,
segments=segment,
cutoff_ranges={"low_limit": (0.15, 0.30), "high_limit": (0.08, 0.20)},
)
 
for name, s in report["segments"].items():
decisions, bands = s["decisions"], s["bands"]
low, high = s["cutoff_range"]
print(f"{name} n={s['n']:,} bad rate {s['bad_rate']:.1%} cutoff range {low:.0%}-{high:.0%}")
print(f" Gini teacher {s['teacher_gini']:.3f} -> artifact {s['gini']:.3f} ({s['gini_retention_pct']:.1f}% retained)")
print(f" decided differently from the teacher: at most {decisions['max_disagreement_rate']:.1%}, "
f"{decisions['mean_disagreement_rate']:.1%} on average across the range")
print(f" approved bad rate, artifact minus teacher: at most {decisions['max_bad_rate_gap']:+.2%}")
print(f" band edges inside the range: {bands['edges_in_range']} (cutoff expressible: {bands['cutoff_expressible']})")
print()
 
print("worst retention :", report["worst_retention_segment"])
print("worst disagreement:", report["worst_disagreement_segment"])

Output#

Captured from an actual run against compileml 0.9.0 and the UCI credit panel. If this script stops working, the build fails.

captured in CIretention_by_segment.py
high_limit n=5,574 bad rate 19.2% cutoff range 8%-20%
Gini teacher 0.541 -> artifact 0.534 (98.8% retained)
decided differently from the teacher: at most 16.2%, 10.8% on average across the range
approved bad rate, artifact minus teacher: at most +0.50%
band edges inside the range: 5 (cutoff expressible: True)
 
low_limit n=1,926 bad rate 30.7% cutoff range 15%-30%
Gini teacher 0.502 -> artifact 0.491 (97.8% retained)
decided differently from the teacher: at most 13.0%, 8.0% on average across the range
approved bad rate, artifact minus teacher: at most +1.27%
band edges inside the range: 2 (cutoff expressible: True)
 
worst retention : low_limit
worst disagreement: high_limit

Notes#

  • Disagreement compares the artifact with the teacher at equal approval volume, so it measures ranking rather than two different calibrations.
  • A cutoff on a band ladder can only sit on a band edge. When no edge falls inside a range, cutoff_expressible is False, however well the model ranks there.
  • When one segment pays, tell a budget gap from a structural one before reaching for sample_weight: the tuning guide sets out the levers and what each costs.

API reference: retention_by_segment →