About CompileML

Where it started

CompileML started in credit risk, where two schools of thought meet and rarely agree. Risk practitioners build scorecards: transparent, reproducible, deployable anywhere, and trusted because a validator can check every point by hand. Data scientists build tree ensembles, which predict better and arrive with a Python environment, a serving stack, post-hoc explanations, and a model no validator can reproduce independently.

Each side is right about what the other gives up, which leaves teams choosing between the model they can defend and the model that performs. The choice is not unique to credit. It appears wherever a model makes a decision about an individual case and has to answer for it — to a regulator, an auditor, a customer, a clinician — or has to run somewhere the data-science stack does not.

CompileML removes that choice: train the strongest teacher you can, then compile it into integer decision logic that reproduces exactly, explains itself by arithmetic, and runs where the decision already happens. Credit remains the origin and the worked example throughout these docs. Where it fits covers the other domains, and where the fit is weaker.

How the project is run

What a model-risk team needs in order to assess CompileML as a dependency. Each CI fact below is read from the library's own configuration at the release these docs describe.

Maintainer
One named maintainer, Carlos Ortiz (@orgoca), with contributions from outside the project — see the contributors.
Every change
Gated by CI before it merges: the test suite on Linux, macOS, Windows across Python 3.10–3.13; the same seeded artifact built on Linux, macOS, Windows with byte-identical hashes required; generated COBOL compiled and run with GnuCOBOL, and generated SQL executed in SQLite, both checked against the Python runtime; and a strict documentation build. CI runs.
Versioning
Semantic versioning, with every change recorded in the changelog.
Releases
Every release is archived on Zenodo. The concept DOI 10.5281/zenodo.22242231 resolves to the latest version and lists each one; CITATION.cff in the repository gives the citation.
Licence
Apache-2.0, provided as is. The library certifies no model, and no regulatory framework in any domain.
Conduct
Everyone taking part follows the Code of Conduct.
Questions
GitHub Discussions — Q&A for how-to questions, Ideas for open design questions, Show and tell for what compiling cost on your data, and Announcements. Issues are for bugs and scoped work. Whether segmented scorecards should ship as one artifact is an open question in Ideas, not a planned feature.