Examples · Selecting

score_batch

Score a whole matrix through the artifact with NumPy, and confirm it produces the runtime's integers exactly.

Code#

score_batch.py
import numpy as np
from _setup import X_test, artifact
 
from compileml.batch import score_batch
from compileml.runtime import decide
 
batch = score_batch(artifact, X_test)
rows = [decide(artifact, row.tolist(), explain=False) for row in X_test]
 
print(f"rows scored : {len(X_test):,}")
print(f"keys : {', '.join(batch)}")
print()
 
# The batch scorer is learning-side convenience — a population scored at once
# for selection, monitoring or a fairness cut — and not a second
# implementation of the decision. It has to agree with the runtime on every
# governed integer, on every row.
for key in ("latent_int", "band_idx", "pd_ppm"):
same = np.array_equal(batch[key], np.array([row[key] for row in rows]))
print(f"identical {key:<10}: {same}")
 
print()
print("first three rows:")
for i in range(3):
print(f" latent_int {batch['latent_int'][i]:>4} band_idx {batch['band_idx'][i]} pd {batch['pd'][i]:.6f}")

Output#

Captured from an actual run against compileml 0.9.0 and the UCI credit panel. If this script stops working, the build fails.

captured in CIscore_batch.py
rows scored : 7,500
keys : raw_micro, latent_micro, latent_int, band_idx, pd_ppm, pd
 
identical latent_int: True
identical band_idx : True
identical pd_ppm : True
 
first three rows:
latent_int 341 band_idx 8 pd 0.378049
latent_int 322 band_idx 8 pd 0.343195
latent_int 279 band_idx 7 pd 0.285714

Notes#

  • Learning-side only. The deployed path is still the standard-library runtime, or the SQL and COBOL exports.
  • It is a faster route to the same integers, not a second implementation of the decision: the example compares every governed integer against decide() row by row.

API reference: score_batch →