Measured runs

What the trade has looked like

Compiling is a trade, and each run here records both sides of it: what the artifact gave up against its teacher, and what it gained — decision latency, artifact size, and whether the hash reproduced. A record rather than a benchmark, and it grows as people publish theirs. Today it holds two, both of them ours, which is the weakest version of this page it will ever be. A hundred runs across other people's data would answer the question; two of our own can only illustrate it.

It will be biased, and the bias will not go away. Nobody reports a disappointment. A compilation that gave up eight points of Gini is exactly the entry this record most needs and is least likely to receive, so read what follows as the favourable edge of what compilation costs rather than the middle of it. If yours went badly, that is the more useful number.

With enough entries it can be split by what the model was for — credit risk, fraud, churn — because retention on a 5/95 fraud panel is a different question from retention on a bureau scorecard, and averaging the two answers neither. Reporting a run takes one JSON block and a paragraph of context: shape, base rate, teacher family, outcome. No institution, no feature names, no data.

The seeded run

Deterministic synthetic credit data — 40,000 rows, 23 features, seed 42 — through the pure-Python runtime. Every number is read from the results file the run commits, not transcribed.

The rows are split three ways, stratified: Fit (24,000) trains everything, including calibration and band edges; Select (8,000) chooses the target, the tree count and the depth from 40 configurations; and Report (8,000) is read once, for the figures below. A holdout that both chooses a configuration and reports its retention flatters itself, which is the reason for the third partition. The winner — α = 0.75, 80 trees, depth 2 — was refit on Fit and Select together before Report was read.

Retention is the artifact's Gini over the ceiling's on Report. The ceiling is the strongest model trained under a declared budget — here a HistGradientBoostingClassifier — and it never ships; retention is measured against it, so it means only as much as that budget does. The floor is reported at equal weight: retention prices the guarantees, and the ratio against a weight-of-evidence logistic regression says whether adopting them was worth it.

Measured on compileml 0.9.0, Python 3.14.3, Windows 11, Intel64 Family 6 Model 170 Stepping 4, GenuineIntel. Latency is measured in a fresh process, after a 30 s rest following selection, and the artifact timed below is the one this run selected, at 80 trees — both of which move the figures, so read them against that description rather than against an earlier release's.

Read them for their order of magnitude, not their third digit. Latency and artifact size barely move with the data: they are properties of the runtime and of the compiled document. Retention is not. It depends on the dataset, the teacher and the depth budget, and a single figure quoted to a decimal place would imply a precision no dataset can promise for another.

reproduce
python benchmarks/run_benchmarks.py
MetricValue
Ceiling Gini0.664
Compiled integer artifact Gini0.651 — 98.0% retained (97.4–98.7)
Floor Gini, WoE logistic regression0.630 — artifact at 103.4% (102.5–104.2)
Band-ordinal Gini0.643 — 96.7% retained
Spearman correlation, ceiling vs. artifact0.980
Selected on Select, of 40 configurationsα = 0.75, 80 trees, depth 2
Score + band + calibrated PD0.021 ms median
Score + band + calibrated PD, p950.025 ms
Full explained decision, 80 trees0.416 ms median
Full explained decision, p950.679 ms
Band assignment alone0.2 µs
Artifact size70 KB
Same configuration and identical hash on rerunYes

What the other configurations scored

The run scores every configuration on Select and commits the whole curve beside its results. 4 of 40 finished within one standard error of the best, and among those the simplest wins: labels over soft targets, then fewer trees, then lower depth. The ten strongest, by Select Gini:

αTreesDepthSelect Gini 
0.7516020.6679tied
0.516020.6670tied
0.58020.6594tied
0.758020.6587selected
116020.6568
016020.6546
18020.6537
0.2516020.6537
0.258020.6500
08020.6479

Select-partition figures, one row per configuration; they rank configurations and do not describe the shipped artifact.

Which is what the third partition is for: the configuration at the top of this table is not the artifact, and its Select figure is not the retention. The selected configuration scored 0.659 here, and the artifact refit from it scored 0.651 on Report.

The public panel

The run above is synthetic — data this project generated for itself, which is the weakest kind of evidence a project can offer about its own results. This is the same protocol on the public UCI credit-card default panel, from the compile_selected example, which executes on every build:

captured in CIcompile_selected.py
selected : alpha=1, 20 trees, depth 2
ceiling Gini 0.536
artifact Gini 0.520 KS 0.399 Brier 0.138
floor Gini 0.514 (WoE logistic regression)
retention 96.9% (95% CI 94.7–99.1)
vs the floor 101.1% (95% CI 99.2–103.2)

A harder problem, a weaker ceiling, and a different answer. On this panel the selection chose labels alone over any blend with the ceiling's predictions, and took the smallest model in a tie band of 27. Retention is lower than the seeded run's, and the interval against the floor reaches below 100%: here, compiling bought the guarantees for roughly what a logistic regression already scored. That is the kind of answer this protocol exists to produce, and it is worth more than a headline that only ever goes one way.

What explanation costs

Scoring is very fast; a fully explained decision costs more — 0.416 ms median in the table above, for the score, band, probability, exact attribution and reason codes together.

The exact pairwise attribution is aggregated per tree. A depth-2 tree is walked at most 8 times however wide the model is, so the cost follows the number of trees rather than the number of features. A perturbation-based derivation of the same integers instead scores the whole ensemble

2 + p + p(p−1)/2

times for p features. Both were timed on the same 120-tree ensemble, attribution alone, median of 40 rows, and reach identical integers on every row timed:

FeaturesTree walks, per treeTree walks, perturbationms, per treems, perturbation
87924,5600.6020.935
2391233,3600.6926.87
50944153,2400.73634.007
10088 split on936606,2400.744133.883

The walk counts are exact and hold on any machine; the milliseconds belong to the one named above. Cost follows walks, and walks follow tree structure rather than width. The models behind this table are fitted to a target that uses every feature, so more of their trees split on three distinct features than the headline model's do — which is why attribution alone at 23 features here costs slightly more than the full decision in the first table.

In practice: explain everything. At 0.416 ms a fully explained decision is real-time for most decisioning — a bureau or data lookup usually costs more — and complete attribution on every decision turns portfolio questions such as driver drift or a fairness cut into census facts rather than sample estimates. It also means every production decision carries its own explanation in the record, computed at decision time under the same artifact hash.