make_credit_data
from compileml.datasets import make_credit_data
Generate a synthetic credit-style dataset with a known structure.
Parameters#
| Name | Type | Default | Kind |
|---|---|---|---|
| n_rowsNumber of rows. | int | 40000 | positional |
| n_featuresNumber of features. Only the first five carry signal; the rest are standard normal noise, which is realistic and keeps the attribution cost formula honest at a defensible feature count. | int | 23 | positional |
seedSeeds a numpy.random.Generator. The same seed gives the same bytes on any platform NumPy supports. | int | 42 | keyword-only |
Returns#
tuple[numpy.ndarray, numpy.ndarray, list[str]]
(X, y, feature_names) — X float64 (n_rows, n_features),
y int, and generic f00… names for the noise columns with
meaningful names for the five that matter.
Raises#
- ValueError
Read from the function body and the private helpers it calls, not inferred. A function that raises nothing has no section here.
Notes#
The committed benchmark in benchmarks/run_benchmarks.py keeps its
own copy of this process on purpose: it draws from a generator that is
shared with later timing code, so routing it through this function
would shift the random stream and move every published benchmark
number. Do not consolidate them without regenerating the benchmark.