compileml.tune

sweep_whitebox

def sweep_whitebox(X, teacher_latent, y, *, trees_grid=(10, 20, 40, 80, 160), depth_grid=(1, 2, 3), alpha_grid=None, learning_rate: float = 0.2, random_state: int = 42, X_val=None, y_val=None, teacher_latent_val=None, monotone_constraints=None, reference=None, explain_timing_rows: int = 3, segments=None, sample_weight=None) -> list[dict]

from compileml.tune import sweep_whitebox

Grid-sweep whitebox capacity; measure what each configuration buys.

Per configuration: Gini and retention vs the teacher, Spearman rank agreement with the teacher, whether attribution is exact (depth <= 2), the quantized model's JSON size, and a measured per-row exact-explanation cost.

Supply a holdout (X_val / y_val / teacher_latent_val) for honest numbers; in-sample retention flatters every configuration.

Pass monotone_constraints to sweep the constrained backend instead — run both and diff the retention column to measure the monotonicity premium before committing to it.

alpha_grid sweeps the *target*: each configuration trains on alpha * y + (1 - alpha) * teacher_latent, so 0.0 is pure distillation (the default, and the historical behavior) and 1.0 trains directly on labels. Whether an interior blend wins is a question about your data rather than a settled one — soft targets sometimes regularize, and at a starved capacity they sometimes just fit teacher noise. Keep the axis orthogonal to capacity by sweeping trees and depth alongside it; a target effect measured at one capacity is easily a capacity effect in disguise. Pass teacher_latent=None to drop the teacher entirely, which forces alpha_grid=(1.0,).

reference accepts a fitted :class:`~compileml.reference.ReferenceModel` or a bare Gini float, and adds reference_gini / gini_vs_reference_pct / beats_reference to every row. Reporting teacher retention without it warns: a ceiling alone cannot tell you the whitebox is losing to a logistic regression.

segments — a label per *evaluation* row (the holdout when one is given) — adds per-segment retention to every row, plus worst_segment and worst_segment_retention_pct, so a portfolio average cannot hide a segment that paid for the compression. Decision agreement across a segment's cutoff range needs a calibrated, banded artifact, so it lives in :func:`~compileml.tune.retention_by_segment` rather than here.

sample_weight (one per training row) is passed to every fit, so a weighted sweep beside an unweighted one shows what moving the budget buys one segment and costs the others.

Parameters#

NameTypeDefaultKind
X—requiredpositional
teacher_latent—requiredpositional
y—requiredpositional
trees_grid—(10, 20, 40, 80, 160)keyword-only
depth_grid—(1, 2, 3)keyword-only
alpha_grid—Nonekeyword-only
learning_ratefloat0.2keyword-only
random_stateint42keyword-only
X_val—Nonekeyword-only
y_val—Nonekeyword-only
teacher_latent_val—Nonekeyword-only
monotone_constraints—Nonekeyword-only
reference—Nonekeyword-only
explain_timing_rowsint3keyword-only
segments—Nonekeyword-only
sample_weight—Nonekeyword-only

Returns#

list[dict]

Raises#

  • ValueError

Read from the function body and the private helpers it calls, not inferred. A function that raises nothing has no section here.

Example#

python
from compileml.tune import sweep_whitebox
 
results = sweep_whitebox(
X_train,
teacher_scores,
y_train,
alpha_grid=(0.0, 0.5, 1.0),
reference=ref,
)

Worked example: sweep_whitebox →