FAQ
How many trees should the whitebox have? What depth?#
Measured answers beat rules of thumb: run
sweep_whitebox on a holdout and pick the elbow of the
retention curve. The short version: trees are a cheap linear knob that never
costs you a guarantee; depth is the knob that takes exact attribution away
above 2. Spend on trees, be stingy with depth.
Why is depth 2 the default?#
It is the boundary of two guarantees at once. At depth ≤ 2 the pairwise attribution is complete — the residual is zero as an arithmetic identity — and the artifact collapses into an exact scorecard. At depth 3 both break, for the same mathematical reason: three-way structure appears. The tooling warns, records, draws, and refuses accordingly.
If I keep adding capacity, don't I just get the teacher back and lose the point of compiling?#
No. Fidelity converges toward the teacher; the compilation properties —
integer determinism, one hashed artifact, SQL/COBOL export, sub-millisecond
scoring — hold at any size. Only exact attribution is at risk, and only from
depth. What grows with trees is artifact bytes and explanation milliseconds,
linearly, and sweep_whitebox prices both.
How many bands should I use?#
Either let the data answer — semantic_bands / governance_bands
return the band count they can statistically defend, and honestly return
one band on noise — or sweep fixed K with sweep_bands and read the
retention-vs-K table.
How do I know my banding isn't leaving money on the table?#
band_efficiency. The
gini_gap is the discrimination your ladder discards; per-band within-band
AUCs (with bootstrap CIs) tell you where — a band whose CI sits above 0.55
can still rank risk internally and is a refinement candidate. Validation
check 4 carries the same numbers on every run.
Can I get a classic points scorecard out of this?#
Yes — exactly, not approximately, at depth ≤ 2: build_scorecard(artifact)
or compileml scorecard decision.json --format csv. The printed tables
re-sum to every production decision bit-for-bit; a validator can reproduce
scores in a spreadsheet.
Can I force a direction — "more delinquency must never score better"?#
Yes. Pass monotone_constraints (per-feature −1/0/+1) to train_whitebox
and the whitebox is trained with scikit-learn's histogram GBM, which enforces
directions during growth. Pass the same declaration to build_artifact and
it is re-verified against the compiled integer trees — the build refuses on
any violation, whatever trainer produced the model — then recorded in the
artifact under the hash. Validation check 9 repeats the verification from the
artifact alone, and at depth ≤ 2 the printed scorecard certifies the
aggregate direction (scorecard_monotone_report). Constraints cost fidelity
wherever the teacher genuinely wiggles; measure the
premium instead of
guessing it.
My whitebox keeps 95% of the teacher. Is that good?#
It is half an answer. Retention measures distance to a ceiling and cannot
tell you the artifact is being beaten by a logistic regression on the same
data — which does happen, particularly at depth 2 where capacity is the
binding constraint. Fit the other side of the comparison with
compileml.reference.fit_reference (or pass your champion scorecard's Gini
as a plain float) and every sweep row carries the floor beside the ceiling;
compile_selected reports both at equal weight, with intervals, and refuses
to build an artifact that does not clear the floor. Validation check 10
applies the same comparison to a built artifact. See
the tuning guide.
Should I distil from a teacher, or train on labels directly?#
Let the pipeline select it. compile_selected trains candidates on the
labels, on the ceiling's cross-fitted predictions, and on blends of the two,
and picks on rows it will not report on. On a 2.9M-row credit portfolio the
whitebox trained on labels beat the distilled one, so distillation is no
longer the default anywhere; soft targets earn their place when events are
scarce and the whitebox would otherwise fit label noise. Sweeping by hand is
still possible with sweep_whitebox(alpha_grid=...) — sweep the target
alongside capacity, since the two are easy to confuse, and use out-of-fold
ceiling predictions, since in-sample ones partly restate the labels.
Why not just use SHAP?#
TreeSHAP is exact for trees and a fine analysis tool — the differences are
about deployment, not correctness. CompileML's explanation is computed on
the deployed object itself (not the pre-compilation model), in integer units
that re-sum to the decision, by a runtime with no ML dependencies. The
explanation is part of the decision record, under the artifact's hash, rather
than a separate analysis run that must be trusted to match. With
explain=True, the SQL and COBOL exports carry the same reason codes and
integer impacts into the warehouse and onto the mainframe.
Why not PMML, ONNX, or m2cgen?#
PMML, ONNX and m2cgen solve related but different problems, and each is the better choice when its problem is yours. PMML is mature, widely understood by validators, and its Scorecard model already carries points and reason codes. ONNX is the broad inference standard, and the natural choice for a model such as a neural network you do not want to compile through a whitebox. m2cgen turns a trained model into readable code in many languages, including ones CompileML does not export to.
All three run the model you trained. CompileML runs a compiled whitebox — a depth-2 whitebox that gives up some discrimination, about 2% of Gini on the committed benchmark — in exchange for a different guarantee. It targets the decision, not just the scorer: integer arithmetic fixed in the artifact rather than left to each evaluator's floating point, so the Python runtime, SQL and COBOL agree to the integer; calibration, bands and reason codes in one hashed document; and attributions for the whole tree ensemble that add back to the score exactly. If you need to run a model elsewhere, use one of those tools. If you need the deployed decision itself to be deterministic and auditable, and can afford the compression, that is the problem CompileML is designed to solve.
Can I compile my XGBoost classifier directly?#
Directly compiled models must emit a latent in [0, 1] — a classifier's raw
margin lives in log-odds space and will be clamped into nonsense. Train a
whitebox on its probabilities instead:
train_whitebox(X, model.predict_proba(X)[:, 1]). Regressors on
probability-like targets compile directly, and the build warns when sample
latents fall outside range.
What about neural networks?#
Same route: any model that produces a probability-like latent can teach a whitebox. The artifact never contains the network — it contains the compiled trees, with the retention measured and recorded.
What happens when I retrain or recalibrate?#
Two mechanically distinct cases, distinguishable by hash. Recalibration
(recalibrate_artifact) refits the PD table on fresh outcomes while the
model and band edges stay byte-identical — provably zero band churn, with the
predecessor's hash recorded as a provenance chain. Retraining produces a
genuinely new artifact and a fresh governance cycle. See
zero-churn recalibration.
How are missing values handled?#
By declared policy inside the artifact: "baseline" re-applies the
training-time imputation at decision time; "reject" refuses the row.
NaN never routes silently through a tree comparison, in any runtime.
Is the artifact hash a signature?#
No — it is an integrity check. Loaders verify it by default and refuse a tampered or corrupted document, but it does not prove who produced the artifact. Provenance of authorship is your repository's and your process's job.
My semantic_bands returned one band. Is that a bug?#
It is the honest answer: under your eps_auc strictness, the data cannot
statistically support discrete classes — either outcomes are too noisy or
the within-band separation you demanded isn't there. Loosen eps_auc
deliberately, or accept that band boundaries would be arbitrary. A banding
tool that always returns the requested K is a random number generator with
labels.
Why is the full explanation slower than scoring?#
It is still more work than scoring, but no longer dramatically so, and no
longer quadratic in feature count. Attribution is aggregated per tree, which
makes the cost O(trees) and independent of p — on the committed
benchmark's 120-tree ensemble, 0.6–0.7 ms per row whether the model has
8 features or 100. More trees cost proportionally more; more features do not.
That is real-time for credit decisioning, which is why explain everything is the recommended default, and it is what makes full-book batch re-explanation practical rather than a 23-CPU-hour job.
Same artifact, same input — could two machines ever disagree?#
Not within the contract: scoring is integer addition, banding integer
comparison, calibration integer lookup, and the single float operation
(x <= threshold) is exact under IEEE 754. CI proves it continuously —
committed reference integers replayed on three OSes, generated SQL executed
and diffed row-for-row, generated COBOL compiled and run. What is not
claimed: that your upstream feature pipeline produces identical bytes across
systems (precisely stated).