Examples · Compiling

train_whitebox

Fit a shallow whitebox against a target and read the fidelity metrics that say whether it is close enough to the ceiling to compile.

Code#

train_whitebox.py
from _setup import X_train, teacher_scores
 
from compileml.compile import train_whitebox
 
# teacher_scores are HistGradientBoostingClassifier probabilities on X_train.
whitebox, fidelity = train_whitebox(X_train, teacher_scores)
 
# Rounded on purpose. These are float training metrics, not governed
# outputs: the last digit of a Pearson correlation differs between BLAS
# builds, and CompileML's determinism claim is about the integer artifact.
for name, value in fidelity.items():
print(f"{name:<24} {value:.6f}")

Output#

Captured from an actual run against compileml 0.9.0 and the UCI credit panel. If this script stops working, the build fails.

captured in CItrain_whitebox.py
pearson 0.967098
spearman 0.921699
mae 0.036113
rmse 0.050300
min_pred 0.047540
max_pred 0.804729

Notes#

  • compile_selected chooses the target, the tree count and the depth for you, and reports on rows that chose nothing. This is the same fit, one configuration at a time.
  • Spearman is the metric that matters: banding is a ranking problem, so rank agreement with the ceiling predicts how much Gini survives compilation.
  • A depth of 2 or less keeps pairwise attribution exact. Deeper whiteboxes compile fine but record exact_attribution: false in the artifact.

API reference: train_whitebox →