Tuning the compilation

Every question on this page has a measured answer — compileml.tune exists so configuration is a table you read, not a guess you defend.

The two knobs are not symmetric#

KnobBuysCostsDeterminism / portability
n_estimators ↑fidelity to the teacherartifact size and explanation time, both linearunaffected
max_depth ↑fidelity per treeexact attribution above 2, per-tree explain cost, scorecard legibilityunaffected

Integer exactness, cross-runtime determinism, and hash governance are structural: a 500-tree artifact is exactly as deterministic and exactly as COBOL-exportable as a 30-tree one. Nothing about compilation degrades with size — what degrades above depth 2 is explainability, and only that.

Spend on trees; be stingy with depth.

Let the data choose: compile_selected#

Two things used to be assumed. That the whitebox should learn from the teacher's probabilities — and on a 2.9M-row credit portfolio the whitebox trained on labels beat the distilled one, with bootstrap intervals excluding zero, and a WoE logistic regression beat both compiled models (#39). And that one holdout could both choose a configuration and report its retention — which leaks: the figure chosen on a holdout is optimistic on it.

compile_selected replaces both assumptions with a protocol:

from compileml import compile_selected

result = compile_selected(
    X, y,
    ceiling=lambda X_fit, y_fit: XGBClassifier(**tuned).fit(X_fit, y_fit),
    ceiling_budget={"trials": 200, "cv": "5-fold on fit"},   # recorded, so retention is comparable
    reference="woe",                                          # or your champion scorecard's Gini
    date=application_date,                                    # Report becomes the latest slice
    group=applicant_id,                                       # one applicant, one partition
)
result.selection_curve    # every configuration, scored on Select
result.selected           # alpha, trees, depth, and the tie band it won inside
report = result.report()  # Report, read once

Three roles. The ceiling — the teacher — is the strongest model you can train under a declared budget. It prices compilation, it may supply soft targets, and it never ships. The floor is a WoE logistic regression fitted inside Fit, or a champion scorecard's Gini. The candidate is the depth-≤2 whitebox, one per configuration.

Three partitions. Fit trains everything, including calibration and band edges. Select chooses the target, the tree count, the depth and the monotone set. Report is read once, for the published figure. The default split is 60/20/20 stratified; date= makes Report the latest slice, and group= keeps each applicant in one partition, so a repeat applicant cannot sit in Fit and Report at once.

Soft targets are cross-fitted. A ceiling scoring its own training rows has partly memorised their labels, so in-sample predictions restate the label and dilute the soft-target arm. The factory is called once, its configuration is frozen, and that configuration is refit K times for out-of-fold predictions: one search plus K fits, never K searches.

The candidate is scored as an artifact. Every configuration is quantized, calibrated and banded on Fit and scored on Select as an integer artifact, so quantization and banding are priced, not just the tree fit. The tree axis is free: the first 20 trees of a 160-tree fit are the 20-tree model, so one fit per (α, depth) covers every tree count.

One-SE-simplest. The best configuration's Select metric is bootstrapped for its standard error; everything within one SE is tied; among the tied, α = 1 wins, then fewer trees, then lower depth. The α preference is a governance prior — an artifact trained on labels has no training dependency on the ceiling — and provenance records it as the tie rule. On large data the SE is tiny and the data decides; on small data the prior decides.

The floor gates. If the selected configuration does not beat the floor on Select, you get a result and no artifact, unless allow_below_floor=True. Retention against the ceiling and the ratio against the floor are reported at equal weight: the first prices the guarantees, the second says whether adopting them is worth it.

The winner is refit on Fit ∪ Select before Report is read. The selection curve therefore ranks configurations; the Report figure is the only number that describes the artifact that ships. result.report() writes that figure into the artifact's hash-covered provenance block and refuses to run twice — hygiene, not proof, and it says so.

Small data. Below min_select_events (300) in Select, selection runs as nested cross-validation over Fit ∪ Select, with a warning; Report stays held out.

Events per split. Every configuration reports eps, the positives in Fit per split in the ensemble. Soft targets help when events are scarce and the whitebox would otherwise fit label noise; the boundary measured so far — soft targets winning below roughly 10 events per split — came from one in-sample sweep, is a lower bound, and is being replaced by #76. It is a diagnostic, never a selection input.

Cost. The search uses the histogram backend, which at 300k rows fitted in a second where the classic backend took five minutes; the tree axis costs nothing; Select is scored with the NumPy batch scorer. The ceiling's refits dominate: budget for about K + 2 fits of your strongest model beyond the one you would train anyway.

Finding the tree count by hand#

from compileml.tune import sweep_whitebox

rows = sweep_whitebox(
    X_train, teacher_latent_train, y_train,
    trees_grid=(20, 40, 80, 160, 320),
    depth_grid=(1, 2),
    alpha_grid=(0.0, 0.5, 1.0),         # the target is a choice; say which
    X_val=X_val, y_val=y_val, teacher_latent_val=teacher_latent_val,
)
# pandas.DataFrame(rows) if you like tables

Each row reports holdout Gini and retention versus the teacher, Spearman rank agreement, exact_attribution, the quantized model's JSON size, and a measured per-row exact-explanation cost. Read it like a cost curve: retention climbs steeply, then plateaus; pick the elbow. In the repository benchmark, the selected configuration — α = 0.75, 80 trees, depth 2 — retained 98.04% of a 300-tree ceiling's Gini on Report (95% interval 97.38–98.72%); 160 trees scored basis points higher on Select and lost the tie to the smaller model. Beyond the plateau you pay linear size and explain time for basis points of fidelity.

A holdout (X_val=) is better than in-sample, where retention flatters every configuration — but a holdout used both to choose a configuration and to report its retention leaks, and sweep_whitebox now warns when it sees one. Choose on it; report on rows never used to choose. That is what compile_selected does for you.

Choosing depth: one table, one story#

depthAttributionScorecardFidelity
1exact, zero residualexact classic scorecard (bin → points)lowest
2 (default)exact, zero residualscorecard + explicit interaction gridsgood
3+residual appearsnone existshighest

Depth ≤ 2 is not a style preference — it is the boundary of two guarantees. The pairwise decomposition is complete for functions with no three-way interactions, which is precisely what depth ≤ 2 trees are; at depth 3 the leftover becomes a real, reported residual (why), and no clean scorecard exists for the same mathematical reason. The tooling enforces the boundary honestly rather than hiding it:

  • train_whitebox warns at max_depth > 2;
  • the artifact records whitebox_max_depth and exact_attribution;
  • every explained decision reports its residual, and the waterfall draws it as an explicit bar;
  • decide() refuses a nonzero residual on an artifact claiming exactness;
  • build_scorecard raises above depth 2.

To quantify what depth 3 would cost you before committing, explain a sample and look at the residuals directly:

residuals = [
    abs(decide(artifact3, row, include_contributions=True)
        ["attribution_residual_half_micro"]) / (2 * 1_000_000)
    for row in X_sample
]
# share of decisions with any unexplained remainder, and how large it gets

"Doesn't a bigger whitebox just become the teacher?"#

Only in the sense you want. More capacity converges toward the teacher's predictions while every compilation property — integer determinism, the hashed single artifact, SQL/COBOL export, sub-millisecond scoring — holds at any size. The one thing you can lose is exact attribution, and that is controlled solely by depth, not trees. The trade to actually manage is pragmatic: linear growth in artifact bytes and explanation milliseconds, which sweep_whitebox prices per configuration.

How many bands?#

Two philosophies, both shipped:

Discover K. semantic_bands and governance_bands return the number of bands the data can statistically defend — Jeffreys-CI separation between neighbors, no residual rank power within any band, bootstrap-certified. Feed them noise and they honestly return one band.

Sweep fixed K.

from compileml.tune import sweep_bands

rows = sweep_bands(latent_train, y_train, k_grid=(4, 6, 8, 10, 12, 16))

Per K: band-ordinal Gini and retention, the Gini gap, the worst within-band AUC with its verdict, minimum band volume, monotonicity violations, and any integer-edge collisions (a K too fine for the display scale to represent — the build would refuse it anyway).

The ceiling a clipped latent puts on K#

There is an upper bound on K that has nothing to do with discrimination.

train_whitebox fits squared error to the teacher's probabilities, so on a low base rate a share of its predictions come out below zero and clip to exactly zero. Those rows are then pinned: they share one latent value and no cut can separate them. Once the pinned mass is larger than one band's worth of volume — roughly when

share_pinned > 1 / n_bands

— at least one quantile cut lands inside the mass, both its edges round to the same integer at the display scale, and that band would be empty.

The builders handle it: colliding edges are dropped, you get fewer bands than you asked for, and a warning says so. Nothing breaks. But it is worth recognising the symptom, because the honest fix is usually upstream. A large pinned share is the same signal build_artifact reports as share_outside_unit_interval, and it means the distillation is overshooting — fewer rounds or a lower learning rate will often recover the bands you wanted.

At a 5% bad rate, roughly 10% of rows pinned is typical, which caps a clean ladder at about ten bands. At 2-3%, expect to lose one or two more.

Money on the table: within-band AUC#

A band ladder discards rank information by design; the governed question is how much and where:

from compileml.bands import band_efficiency

eff = band_efficiency(latent_val, y_val, artifact)
eff["gini_gap"]        # continuous Gini − band-ordinal Gini: the headline
eff["per_band"]        # n, bad rate, within-band AUC with bootstrap CI, verdict
eff["worst_band"]      # strongest refinement candidate

Reading the per-band verdicts:

  • exhausted — CI upper bound ≤ 0.55: the score is used up inside this band; splitting it further separates noise, not risk.
  • refinable — CI lower bound ≥ 0.55: the latent can still rank outcomes inside the band; a finer cut there would separate risk your policy currently treats as homogeneous. That is money on the table.
  • inconclusive — the interval spans both stories; more volume before concluding anything.

The same diagnostics attach to every validation run: check 4 of validate_artifact reports banding_gini_gap and worst_within_band_auc whenever outcomes are supplied, advisory by default and gateable via max_within_band_auc= when your policy wants a hard limit.

Retention where the decisions are made#

A portfolio retention figure is an average, and distillation loss is almost never uniform. It is also dominated by the easy separations — the obviously good against the obviously bad — while a lending decision is made in the narrow band of risk where the cutoff sits, and that band differs by segment. A 99% average can hide a segment at 74%.

retention_by_segment scores a holdout through the built artifact and reports, per segment, what compilation cost there:

from compileml.tune import retention_by_segment

report = retention_by_segment(
    artifact, teacher_latent_holdout, X_holdout, y_holdout,
    segments=file_depth,                        # a label per row
    cutoff_ranges={"thin_file": (0.10, 0.17),   # PD, per segment
                   "thick_file": (0.02, 0.06)},
)
report["worst_retention_segment"], report["worst_disagreement_segment"]

A cutoff range is not a cutpoint. A cutpoint needs its own study before deployment; a range of PD — the risk appetite a segment's cutoff will land in — is enough to establish whether the artifact is sound anywhere that study could put it. Ranges are policy, so they are inputs to the report and never part of the hashed artifact.

For each segment the report carries three things.

Retention — teacher and whitebox Gini, gini_retention_pct and Spearman agreement on the segment's own rows.

Decision agreement across the range. At every cutoff from the low end of the range to the high end, the artifact approves applicants whose emitted PD is at or below it, and the teacher approves the same number of its own lowest-risk applicants. Equal volume keeps the comparison about ranking rather than two different calibrations. Each point records the approval rate, the share of applicants decided differently (disagreement_rate), and the bad rate each model approves; the summary keeps the worst and the average. It reads as: wherever the thin-file cutoff lands between 10% and 17%, at most this share of applicants are decided differently than the teacher would have decided them, and the approved book's bad rate differs by at most this much.

Band resolution. A cutoff on a band ladder can only sit on a band edge. bands.edges_in_range counts the ladder's edges whose calibrated PD falls inside the range; when it is zero, cutoff_expressible is False and no cutoff inside the range exists on this ladder, however well the model ranks. Ladders built for a whole portfolio do this to small segments; rebuild the ladder before anyone starts the cutoff study.

A single (low, high) applies to every segment; a segment without a range gets retention only. Ranges are PD as fractions, so (0.02, 0.06) rather than (2, 6).

sweep_whitebox(..., segments=...) adds per-segment retention and worst_segment to every configuration, so capacity can be chosen by the segment that pays most rather than by the average. The cutoff-range checks need a calibrated, banded artifact, so they run on the one you build.

When one segment pays#

A whitebox is trained to copy the teacher everywhere, on average. It spends its fixed budget where most of the squared error is — the largest segment and the busiest part of the score range — and nothing tells it where a small segment's decisions are made. A segment with low retention has one of three problems, and they need different answers.

CauseHow it showsWhat helps
Budget — too little capacity reaches the segmentits retention climbs as trees_grid growsmore trees, or weighting
Structure — the segment differs in a way depth 2 cannot expressits retention plateaus however many treesits own artifact
Data — few rows or defaults in the segmentthe teacher's own Gini there is weak or unstablecheck the floor, not the teacher

Sweep with segments= to tell budget from structure. The structural case has a precise cause: depth-2 trees express pairwise effects exactly, so "this feature matters differently for this segment" (segment × feature) is reachable, but "these two features interact only in this segment" (segment × A × B) is three-way, and no number of depth-2 trees represents it.

In order of cost:

  1. More trees — helps a budget gap, at a linear cost, with exactness untouched.

  2. Make the segment an input — a segment indicator lets depth 2 express segment-specific effects of single features. It cannot reach a segment-specific interaction.

  3. Weight the fit toward the segment, and if you like toward rows whose teacher probability sits near the segment's cutoff range:

    weights = np.where(segment == "thin_file", 5.0, 1.0)
    near = (segment == "thin_file") & (teacher >= 0.05) & (teacher <= 0.25)
    weights[near] *= 2.0
    model, _ = train_whitebox(X, teacher, n_estimators=40, sample_weight=weights)
    

    How much weight, and how wide "near" is, are modelling judgements, which is why they are not defaults. Weighting moves the budget rather than creating it, so re-check every segment afterwards. It changes how the model ranks, not the PD — calibration is fitted afterwards on outcomes — and the artifact stays exact.

  4. Give the segment its own artifact — the whole budget, its own calibration, and its own band ladder, which also fixes a ladder with no edge in the segment's range. This is the strongest answer to a structural gap, and the one with a governance cost: two artifacts, and routing between them that must itself be auditable (#13).

Do not raise depth to 3 to reach the three-way effect: it gives up exact attribution and the scorecard, and a segment with low retention is often the one receiving the most adverse-action notices.

What the levers did on synthetic data built for each case (a 10% segment, depth 2, holdout retention of the small segment):

GapUnweightedSegment weighted ×5Own artifact
Budget — the segment has its own drivers10.1%87.3%, while the other segment fell from 97.8% to 91.3%—
Structure — a segment-only interaction73.0%75.9%79.3%

Weighting rescued a starved segment and charged the rest for it; it barely moved a structural gap, where a separate artifact did better. Finally, low retention is relative to the teacher. If a plain logistic regression matches the teacher on that segment (reference=), the gap is not compression, and a scorecard base with a residual correction (#39) is the better lever.

Producing a scorecard#

At depth ≤ 2 the artifact is a points-based scorecard — exactly, not as an approximation:

from compileml.scorecard import build_scorecard, scorecard_to_markdown

scorecard = build_scorecard(artifact)
print(scorecard_to_markdown(scorecard, labels=DISPLAY_NAMES))
compileml scorecard decision.json --format csv --out scorecard.csv

Depth 1 collapses to the classic form — per feature, bin → points. Depth 2 adds explicit pairwise interaction grids over the union of the relevant thresholds. Points are the artifact's own integers, and the identity

base_points + Σ main_effect(x) + Σ interaction(x) == raw_micro

holds bit-for-bit on every row (score_from_scorecard re-derives any decision from the printed tables alone — the test suite asserts it). Hand the CSV to a validator and they can reproduce production scores in a spreadsheet.

Above depth 2, build_scorecard raises instead of approximating — the same boundary as exact attribution, for the same reason.

Expect a depth-2 scorecard to be mostly interaction grids. A tree whose two splits use different features contributes a grid rather than a main effect, and boosting seldom spends both splits on one feature. The fairness notebook's 40-tree model on the 23-feature UCI panel compiles to 55 grids and no main effects; an independent evaluation on fraud data found one main effect among 39 grids. The scorecard is still exact, but it is not the one-table-per-feature card many validation functions expect. If yours does, compile at max_depth=1 — the classic form in the table above — and measure the fidelity it costs with sweep_whitebox before committing to it.

Enforcing monotone directions#

A compiled scorecard with a bin where more delinquency scores better is a scorecard a committee rejects on sight — even when the wiggle is statistically justified. Declare the directions and the whitebox is trained with scikit-learn's HistGradientBoostingRegressor, which enforces them during tree growth:

model, metrics = train_whitebox(
    X, teacher_latent,
    monotone_constraints={"UTIL": +1, "TENURE": -1},   # or a [-1, 0, +1, ...] list
)
artifact = build_artifact(
    model, feature_names, baseline, edges,
    monotone_constraints={"UTIL": +1, "TENURE": -1},
    ...,
)

Name-keyed dicts work at train_whitebox when X is a DataFrame; with bare arrays, key by index. Without constraints, nothing changes — the classic GradientBoostingRegressor path is untouched.

The declaration at build_artifact is not a training-time promise passed along: the builder re-verifies the quantized integer trees against it, tree by tree, and refuses to emit the artifact on any violation — whatever trainer produced the model. The verified signs are recorded in the artifact (model.monotone_constraints, hash-covered), validation check 9 re-verifies them from the artifact alone, and at depth ≤ 2 the scorecard's own tables certify the aggregate direction (scorecard_monotone_report) — a check a validator can repeat in a spreadsheet.

The monotonicity premium, measured#

Constraints cost fidelity wherever the teacher genuinely wiggles, and the two backends also regularize differently (histogram binning, leaf-size defaults), so do not guess the cost — measure it:

rows_free = sweep_whitebox(X, latent, y, alpha_grid=(0.0, 1.0),
                           X_val=Xv, y_val=yv, teacher_latent_val=lv)
rows_mono = sweep_whitebox(X, latent, y, alpha_grid=(0.0, 1.0),
                           X_val=Xv, y_val=yv, teacher_latent_val=lv,
                           monotone_constraints={"UTIL": +1, "TENURE": -1})

Diff the gini_retention_pct column at your chosen configuration. If the premium is small, the teacher's wiggle was noise and the constraint bought committee-credibility for free; if it is large, the teacher has learned a genuinely non-monotone pattern, and that is worth investigating before any constraint is imposed.

The floor: what should the whitebox beat?#

gini_retention_pct answers "how close did we get to the teacher?" It cannot answer "did we beat a logistic regression?", and those are different questions. A whitebox at 95% of a very strong teacher can still be losing to thirty logistic coefficients on the same data — retention alone will never say so, because it only measures distance to a ceiling.

So measure the floor too:

from compileml.reference import fit_reference

reference = fit_reference(X_train, y_train, feature_names=FEATURES)
rows = sweep_whitebox(X, teacher, y, reference=reference, ...)

Every row then carries reference_gini, gini_vs_reference_pct and beats_reference beside the teacher columns. Sweeping with a teacher and no reference warns, because a one-sided report is the failure mode this exists to prevent.

fit_reference is a sanity floor, not a challenger model: shallow supervised binning per feature, weight-of-evidence encoding, logistic regression, and regularization strength chosen by cross-validation with the WOE tables refit inside each fold. If you already have a champion scorecard, its Gini is a better floor than anything fitted here — pass the number directly, since reference= accepts a float:

rows = sweep_whitebox(X, teacher, y, reference=0.8515, ...)

The same argument goes to validate_artifact(reference=...) as check 10, advisory by default and gateable with require_reference_floor=True.

Choosing the target: a selected parameter, not a given#

train_whitebox regresses onto whatever target you hand it: the labels, the ceiling's out-of-fold probabilities, or a blend alpha * y + (1 - alpha) * ceiling. Which one wins is a property of your data, not of the method — on the portfolio above the labels won outright — so compile_selected searches α alongside capacity and picks on Select. Pure distillation, alpha_grid=(0.0,), was the old default of sweep_whitebox; it now warns when the grid is not passed, and the default becomes (0.0, 0.5, 1.0) in 1.0.

By hand, alpha_grid sweeps it:

rows = sweep_whitebox(
    X, teacher, y,
    trees_grid=(20, 40, 80),
    alpha_grid=(0.0, 0.25, 0.5, 0.75, 1.0),   # 0 = distil, 1 = labels
    reference=reference,
)

Two cautions. Keep the axis orthogonal to capacity — a target effect measured at a single starved capacity is easily a capacity effect wearing a disguise, so sweep trees and depth alongside it. And do not assume the endpoints bracket the answer: soft targets sometimes regularize, so an interior blend can win. That is the argument for sweeping rather than asserting either end.

To drop the ceiling entirely, pass teacher_latent=None; alpha_grid is forced to (1.0,) and the teacher columns come back None. Whatever you sweep by hand, feed sweep_whitebox the ceiling's out-of-fold predictions: in-sample ones partly restate the labels and understate what a soft target does.

Defaults, for the impatient#

compile_selected(X, y, ceiling=..., reference="woe") with its default grid — α ∈ {0, 0.25, 0.5, 0.75, 1}, trees {20, 40, 80, 160}, depth {1, 2} — is the measured path, and the repository benchmark runs through it. By hand, train_whitebox(n_estimators=30, max_depth=2) and n_bands=10 are sane starting points. The sweeps are for when "sane" needs to become "measured" — which, in a model governance file, it eventually does.