FAQ

How many trees should the whitebox have? What depth?#

Measured answers beat rules of thumb: run sweep_whitebox on a holdout and pick the elbow of the retention curve. The short version: trees are a cheap linear knob that never costs you a guarantee; depth is the knob that takes exact attribution away above 2. Spend on trees, be stingy with depth.

Why is depth 2 the default?#

It is the boundary of two guarantees at once. At depth ≤ 2 the pairwise attribution is complete — the residual is zero as an arithmetic identity — and the artifact collapses into an exact scorecard. At depth 3 both break, for the same mathematical reason: three-way structure appears. The tooling warns, records, draws, and refuses accordingly.

If I keep adding capacity, don't I just get the teacher back and lose the point of compiling?#

No. Fidelity converges toward the teacher; the compilation properties — integer determinism, one hashed artifact, SQL/COBOL export, sub-millisecond scoring — hold at any size. Only exact attribution is at risk, and only from depth. What grows with trees is artifact bytes and explanation milliseconds, linearly, and sweep_whitebox prices both.

How many bands should I use?#

Either let the data answer — semantic_bands / governance_bands return the band count they can statistically defend, and honestly return one band on noise — or sweep fixed K with sweep_bands and read the retention-vs-K table.

How do I know my banding isn't leaving money on the table?#

band_efficiency. The gini_gap is the discrimination your ladder discards; per-band within-band AUCs (with bootstrap CIs) tell you where — a band whose CI sits above 0.55 can still rank risk internally and is a refinement candidate. Validation check 4 carries the same numbers on every run.

Can I get a classic points scorecard out of this?#

Yes — exactly, not approximately, at depth ≤ 2: build_scorecard(artifact) or compileml scorecard decision.json --format csv. The printed tables re-sum to every production decision bit-for-bit; a validator can reproduce scores in a spreadsheet.

Can I force a direction — "more delinquency must never score better"?#

Yes. Pass monotone_constraints (per-feature −1/0/+1) to train_whitebox and the whitebox is trained with scikit-learn's histogram GBM, which enforces directions during growth. Pass the same declaration to build_artifact and it is re-verified against the compiled integer trees — the build refuses on any violation, whatever trainer produced the model — then recorded in the artifact under the hash. Validation check 9 repeats the verification from the artifact alone, and at depth ≤ 2 the printed scorecard certifies the aggregate direction (scorecard_monotone_report). Constraints cost fidelity wherever the teacher genuinely wiggles; measure the premium instead of guessing it.

My whitebox keeps 95% of the teacher. Is that good?#

It is half an answer. Retention measures distance to a ceiling and cannot tell you the artifact is being beaten by a logistic regression on the same data — which does happen, particularly at depth 2 where capacity is the binding constraint. Fit the other side of the comparison with compileml.reference.fit_reference (or pass your champion scorecard's Gini as a plain float) and every sweep row carries the floor beside the ceiling; compile_selected reports both at equal weight, with intervals, and refuses to build an artifact that does not clear the floor. Validation check 10 applies the same comparison to a built artifact. See the tuning guide.

Should I distil from a teacher, or train on labels directly?#

Let the pipeline select it. compile_selected trains candidates on the labels, on the ceiling's cross-fitted predictions, and on blends of the two, and picks on rows it will not report on. On a 2.9M-row credit portfolio the whitebox trained on labels beat the distilled one, so distillation is no longer the default anywhere; soft targets earn their place when events are scarce and the whitebox would otherwise fit label noise. Sweeping by hand is still possible with sweep_whitebox(alpha_grid=...) — sweep the target alongside capacity, since the two are easy to confuse, and use out-of-fold ceiling predictions, since in-sample ones partly restate the labels.

Why not just use SHAP?#

TreeSHAP is exact for trees and a fine analysis tool — the differences are about deployment, not correctness. CompileML's explanation is computed on the deployed object itself (not the pre-compilation model), in integer units that re-sum to the decision, by a runtime with no ML dependencies. The explanation is part of the decision record, under the artifact's hash, rather than a separate analysis run that must be trusted to match. With explain=True, the SQL and COBOL exports carry the same reason codes and integer impacts into the warehouse and onto the mainframe.

Why not PMML, ONNX, or m2cgen?#

PMML, ONNX and m2cgen solve related but different problems, and each is the better choice when its problem is yours. PMML is mature, widely understood by validators, and its Scorecard model already carries points and reason codes. ONNX is the broad inference standard, and the natural choice for a model such as a neural network you do not want to compile through a whitebox. m2cgen turns a trained model into readable code in many languages, including ones CompileML does not export to.

All three run the model you trained. CompileML runs a compiled whitebox — a depth-2 whitebox that gives up some discrimination, about 2% of Gini on the committed benchmark — in exchange for a different guarantee. It targets the decision, not just the scorer: integer arithmetic fixed in the artifact rather than left to each evaluator's floating point, so the Python runtime, SQL and COBOL agree to the integer; calibration, bands and reason codes in one hashed document; and attributions for the whole tree ensemble that add back to the score exactly. If you need to run a model elsewhere, use one of those tools. If you need the deployed decision itself to be deterministic and auditable, and can afford the compression, that is the problem CompileML is designed to solve.

Can I compile my XGBoost classifier directly?#

Directly compiled models must emit a latent in [0, 1] — a classifier's raw margin lives in log-odds space and will be clamped into nonsense. Train a whitebox on its probabilities instead: train_whitebox(X, model.predict_proba(X)[:, 1]). Regressors on probability-like targets compile directly, and the build warns when sample latents fall outside range.

What about neural networks?#

Same route: any model that produces a probability-like latent can teach a whitebox. The artifact never contains the network — it contains the compiled trees, with the retention measured and recorded.

What happens when I retrain or recalibrate?#

Two mechanically distinct cases, distinguishable by hash. Recalibration (recalibrate_artifact) refits the PD table on fresh outcomes while the model and band edges stay byte-identical — provably zero band churn, with the predecessor's hash recorded as a provenance chain. Retraining produces a genuinely new artifact and a fresh governance cycle. See zero-churn recalibration.

How are missing values handled?#

By declared policy inside the artifact: "baseline" re-applies the training-time imputation at decision time; "reject" refuses the row. NaN never routes silently through a tree comparison, in any runtime.

Is the artifact hash a signature?#

No — it is an integrity check. Loaders verify it by default and refuse a tampered or corrupted document, but it does not prove who produced the artifact. Provenance of authorship is your repository's and your process's job.

My semantic_bands returned one band. Is that a bug?#

It is the honest answer: under your eps_auc strictness, the data cannot statistically support discrete classes — either outcomes are too noisy or the within-band separation you demanded isn't there. Loosen eps_auc deliberately, or accept that band boundaries would be arbitrary. A banding tool that always returns the requested K is a random number generator with labels.

Why is the full explanation slower than scoring?#

It is still more work than scoring, but no longer dramatically so, and no longer quadratic in feature count. Attribution is aggregated per tree, which makes the cost O(trees) and independent of p — on the committed benchmark's 120-tree ensemble, 0.6–0.7 ms per row whether the model has 8 features or 100. More trees cost proportionally more; more features do not.

That is real-time for credit decisioning, which is why explain everything is the recommended default, and it is what makes full-book batch re-explanation practical rather than a 23-CPU-hour job.

Same artifact, same input — could two machines ever disagree?#

Not within the contract: scoring is integer addition, banding integer comparison, calibration integer lookup, and the single float operation (x <= threshold) is exact under IEEE 754. CI proves it continuously — committed reference integers replayed on three OSes, generated SQL executed and diffed row-for-row, generated COBOL compiled and run. What is not claimed: that your upstream feature pipeline produces identical bytes across systems (precisely stated).