retention_by_segment
from compileml.tune import retention_by_segment
What compilation cost, per segment and across each segment's cutoff range.
Score X through the artifact and report, per segment:
- **retention** — teacher and whitebox Gini, gini_retention_pct and
Spearman agreement, on the segment's own rows;
- **decision agreement** (when the segment has a range) — at every cutoff
across the PD range, the share of applicants the artifact decides
differently from the teacher at the same approval volume, and the bad
rate each approves;
- **band resolution** (when the segment has a range) — how many of the
ladder's edges fall inside the range, since a cutoff can only sit on
one.
segments is a label per row; omit it to treat the rows as one
segment. cutoff_ranges is PD as fractions: a mapping of segment label
to (low, high), or a single (low, high) for every segment. A
segment without a range gets retention only. Use holdout rows — in-sample
retention flatters every artifact. Cutoff ranges need a calibrated
artifact: without a calibration table the emitted "PD" is the raw score
rescaled, and a PD range would be measured on the wrong scale.
The report names worst_retention_segment and
worst_disagreement_segment, so an average cannot hide one segment.
Parameters#
| Name | Type | Default | Kind |
|---|---|---|---|
| artifact | dict | required | positional |
| teacher_latent | — | required | positional |
| X | — | required | positional |
| y | — | required | positional |
| segments | — | None | keyword-only |
| cutoff_ranges | — | None | keyword-only |
| n_cutoffs | int | 21 | keyword-only |
Returns#
dict
Raises#
- ValueError
Read from the function body and the private helpers it calls, not inferred. A function that raises nothing has no section here.