compileml.datasets

make_credit_data

def make_credit_data(n_rows: int = 40000, n_features: int = 23, *, seed: int = 42) -> tuple[numpy.ndarray, numpy.ndarray, list[str]]

from compileml.datasets import make_credit_data

Generate a synthetic credit-style dataset with a known structure.

Parameters#

NameTypeDefaultKind
n_rowsNumber of rows.int40000positional
n_featuresNumber of features. Only the first five carry signal; the rest are standard normal noise, which is realistic and keeps the attribution cost formula honest at a defensible feature count.int23positional
seedSeeds a numpy.random.Generator. The same seed gives the same bytes on any platform NumPy supports.int42keyword-only

Returns#

tuple[numpy.ndarray, numpy.ndarray, list[str]]

(X, y, feature_names) — X float64 (n_rows, n_features), y int, and generic f00… names for the noise columns with meaningful names for the five that matter.

Raises#

  • ValueError

Read from the function body and the private helpers it calls, not inferred. A function that raises nothing has no section here.

Notes#

The committed benchmark in benchmarks/run_benchmarks.py keeps its own copy of this process on purpose: it draws from a generator that is shared with later timing code, so routing it through this function would shift the random stream and move every published benchmark number. Do not consolidate them without regenerating the benchmark.