train_whitebox
from compileml.compile import train_whitebox
Fit a whitebox GBM to a target.
target is any per-row vector to regress onto: a teacher's latent
probabilities (distillation), the binary labels themselves, or a blend
of the two. Distillation is a choice here, not an assumption — at a
depth-2 budget, fitting a soft target can spend capacity on teacher
noise instead of outcome, so it is worth sweeping rather than assuming
(sweep_whitebox(alpha_grid=...)).
Returns (model, metrics) where metrics quantifies fidelity to the target on the training data (pearson, spearman, mae, rmse, prediction range).
monotone_constraints takes a per-feature sequence of -1/0/+1 or a
dict keyed by feature index (or by name, when X is a DataFrame
carrying column names). Any nonzero sign switches
the backend to HistGradientBoostingRegressor; None (or all
zeros) keeps the classic GradientBoostingRegressor.
backend overrides that choice: "hist" for
HistGradientBoostingRegressor and "gbr" for the classic
GradientBoostingRegressor. The histogram backend is the one to use
at scale — it bins features and runs multithreaded, and was measured at
1–2 s where the classic backend took 18 s at 30k rows and 293 s at 300k
— but the default stays None (classic unless constrained) because
changing it would change every artifact built with the defaults.
sample_weight — one non-negative weight per row — tells the fit where
fidelity matters. A whitebox has a fixed budget and, unweighted, spends
it where most of the squared error is: the largest segment and the busiest
part of the score range. Up-weighting a segment, or the rows near its
cutoff range, moves that budget; other segments pay for it, so re-check
them with :func:`~compileml.tune.retention_by_segment`. Weighting changes
how the model ranks, not the PD: calibration is fitted afterwards on
outcomes. The returned fidelity metrics stay unweighted. Weighting cannot
create an effect depth 2 cannot express — a segment-only interaction is
three-way — see the tuning guide.
Parameters#
| Name | Type | Default | Kind |
|---|---|---|---|
| X | — | required | positional |
| target | — | None | positional |
| n_estimators | int | 30 | keyword-only |
| max_depth | int | 2 | keyword-only |
| learning_rate | float | 0.2 | keyword-only |
| random_state | int | 42 | keyword-only |
| loss | str | 'squared_error' | keyword-only |
| monotone_constraints | — | None | keyword-only |
| teacher_latent | — | None | keyword-only |
| sample_weight | — | None | keyword-only |
| backend | str | None | None | keyword-only |
Raises#
- TypeError
- ValueError
Read from the function body and the private helpers it calls, not inferred. A function that raises nothing has no section here.
Example#
from compileml.compile import train_whitebox whitebox, fidelity = train_whitebox( X_train, teacher.predict_proba(X_train)[:, 1], monotone_constraints={"utilization": +1, "income": -1},)print(fidelity)