Deterministic decision artifacts

Train the strongest model you can. Ship integers.

Your strongest model sets the ceiling. CompileML searches for the best whitebox it can compile, chooses it on data it will not report on, and packages the decision into one hashed JSON document — score, calibrated probability, risk band, reason codes, attribution. The ceiling is a measuring instrument; it never ships.

Get started
pip install compileml
Python
Python 3.10+
Licence
Apache-2.0
Version
v0.9.0
Score
0.021 ms
Runtime deps
0
Artifact
70 KB
decision.json — what production loadshash verified
band
"G09"
pd
0.378049
latent_int
341
reasons_negative[0].code
"REPAYMENT_STATUS_M1"
reasons_negative[0].impact_int
117
artifact_hash
"a31afc9b…"

One document. Integer addition, integer comparison and table lookup — under a standard-library-only Python runtime, or exported as standalone SQL or COBOL.

Why this exists

CompileML started in credit risk, where two schools of thought meet and rarely agree. Risk practitioners build scorecards: transparent, reproducible, deployable anywhere, and short of predictive power. Data scientists build tree ensembles: stronger, and delivered with a Python environment, a serving stack, post-hoc explanations and a model no validator can reproduce. Each side is right about what the other gives up, which leaves a bad choice between the model you can defend and the model that performs.

The choice is not unique to credit. It appears wherever a model makes a decision about an individual case and has to answer for it — to a regulator, an auditor, a customer, a clinician — or has to run where the data-science stack does not: fraud and anti-money-laundering alerts, insurance underwriting and claims, eligibility screening, clinical decision support. Where it fits says where the fit is strong, where it is weaker, and how to read the credit vocabulary from another domain. CompileML removes the choice.

The four objections

Reluctance to put machine learning behind a decision that must be answered for comes down to four things, and they are all reasonable. Each has a structural answer here rather than a reassurance, and each answer is checked on every build. The evidence comes from credit, where the project began.

Scores drift

A score behind a decision should be a fact, not a distribution over environments. Refreshing a model usually means re-cutting the bands, and a case moving band for reasons that have nothing to do with the case.

Leaves are quantized once, at compile time. Recalibration refits probabilities while the model and the band edges stay byte-identical, so updating a PD table cannot move a single account between bands.

captured in CIrecalibrate_artifact
model unchanged : True
band edges unchanged : True
calibration changed : True

From the recalibrate_artifact example, run against the UCI credit panel on every build.

Explanations do not reconcile

An explanation of a regulated decision should not merely resemble the decision. Post-hoc explainers approximate a model they are not part of, and the numbers do not have to add up to anything.

Attribution is computed from the compiled model in integer units. At whitebox depth two or less the artifact collapses into a points scorecard — exactly, not approximately — and the printed table re-derives the production score bit for bit.

captured in CIbuild_scorecard
from the scorecard table : 341214
from the runtime : 341214
identical : True

From the build_scorecard example, run against the UCI credit panel on every build.

The deployment stack is not the modeling stack

Important decisions run on SQL warehouses, core platforms and mainframes. Requiring the training environment in production is often unrealistic, and rewriting the model by hand for the target system produces a second model nobody reconciles.

The artifact exports to standalone SQL or COBOL — score, band and calibrated probability, and with explain=True the reason codes and their integer impacts. Same integers, no model runtime, running where the data already is. The generated query below was executed in SQLite and compared against the Python runtime.

captured in CIexport_sql
SQLite Python
latent_int 341 341
band G09 G09
identical: True

From the export_sql example, run against the UCI credit panel on every build.

A fairness review audits an approximation

A fairness review has to say why groups are treated differently, not only that they are. When the explanation is a post-hoc estimate, the per-feature account of a group gap is an estimate too, and the reasons examined are a ranking standing in for the reasons people are actually given.

The audit reads the decisions production made. Attribution reconciles to the score in integers, so the group score gap splits by feature with nothing unaccounted for, and reason parity is measured on the reason codes the runtime issued. Here the model was never given sex: the audit still traces the gap to features, and records that the counterfactual test does not apply. It produces evidence for a validator, not a certificate of compliance.

captured in CIfairness_audit
adverse impact ratio (Male / Female): 0.974 — inside the 0.80-1.25 window
LIMIT_BAL 35.3% of the gap
Male movement 721,763,812 contributions 721,763,812 equal True
Female movement 1,032,896,530 contributions 1,032,896,530 equal True
§11 counterfactual: not_an_input

From the fairness_audit example, run against the UCI credit panel on every build.

Decision waterfall: a baseline plus integer feature impacts summing to the score
The same decision, drawn by waterfall_svg() from the payload the runtime returned. The plot cannot disagree with the deployed decision because it is not recomputing it — and the renderer imports nothing outside the standard library.

Compile once, decide anywhere

Compilation happens once, wherever you train. Production loads a JSON document.

compile.py
from compileml import compile_selected
from compileml.artifact import save_artifact
 
result = compile_selected(
X,
y,
# The ceiling: your strongest model, under a budget you declare.
ceiling=lambda X_fit, y_fit: XGBClassifier(**tuned).fit(X_fit, y_fit),
ceiling_budget={"trials": 200, "cv": "5-fold on fit"},
reference="woe", # the floor it has to beat
reasons=REASON_DICTIONARY,
)
 
result.selected # target, trees, depth — chosen on Select
report = result.report() # Report, read once: retention and the floor ratio
 
save_artifact(result.artifact, "decision.json")
production.py
from compileml.runtime import load_artifact, decide
 
artifact = load_artifact("decision.json")
decision = decide(artifact, applicant_row)

Or skip the Python runtime entirely:

shell
compileml export decision.json --target sql --out scorer.sql
compileml export decision.json --target cobol --out scorer.cob

What it costs

A depth-2 compilation gives up some Gini against the ceiling, and that is the price of everything above. Retention is the artifact's Gini over the ceiling's, read once on rows that chose nothing. On the seeded run published here it is 98.0% (97.4–98.7), against a ceiling tuned under a declared budget.

Retention alone cannot say whether that is worth adopting, so the floor is reported beside it: the same artifact scores 103.4% (102.5–104.2) of a weight-of-evidence logistic regression on the same rows. A whitebox can sit at 95% of a very strong ceiling and still lose to thirty logistic coefficients; only the two figures together say what compiling bought. Both are properties of a dataset, a ceiling and a depth budget rather than of this library, so measure them on your own data — compile_selected does it the same way.

What does not move with the data: the artifact scores in 0.021 ms median in pure Python and occupies 70 KB as one JSON document. Those are properties of the runtime and of the compiled document.

A fully explained decision — score, band, probability, exact attribution and reason codes — takes 0.416 ms median, against 0.021 ms to score alone. Attribution is aggregated per tree, so its cost follows the number of trees rather than the number of features, and it is exact rather than sampled. Above whitebox depth 2 no exact scorecard exists and attribution residuals appear; the tools raise rather than approximate.

CompileML is not a trainer — the ceiling comes from XGBoost, LightGBM, scikit-learn or elsewhere. Determinism means the same input values with the same artifact produce the same governed outputs; producing the same input values across systems remains the caller's responsibility. And no library can certify an institution's model, data, policy language or governance process.

Every run measured →

What CI checks

Given the same artifact and the same input values, CompileML produces the same governed integer outputs across supported runtimes. The repository tests that rather than asserting it — 8 checks on every build.

  1. SQL output is executed in SQLite and compared row by row with the Python runtime.
  2. Generated COBOL is compiled and run in CI, then checked against the reference implementation.
  3. The same seeded artifact is built on Linux, macOS and Windows and the hashes are compared.
  4. Committed reference decisions are replayed on every OS and Python version in the matrix.
  5. Attribution is added back to the decision during validation.
  6. Recalibration tests verify that the model and band edges remain unchanged.
  7. Scorecard tables are re-summed against the runtime’s own integers.
  8. The standard-library-only runtime is enforced by inspecting its imports.

The artifact carries a SHA-256 hash. Loaders verify it by default and reject a document whose contents no longer match. That detects modification or corruption; it is an integrity check, not a cryptographic signature of who produced the artifact.

Documentation