Examples · Validation & governance
FairnessAudit
Audit the compiled decisions by a protected attribute the model was never given, and decompose the group score gap by feature exactly.
Code#
fairness_audit.py
from _setup import SEED, X_test, artifact, feature_names, y_testfrom sklearn.model_selection import train_test_split from compileml.datasets import load_credit_defaultfrom compileml.fairness import FairnessAuditfrom compileml.runtime import decide # The artifact was compiled without demographics. Load them separately and# repeat the same split, so each holdout row gets its own SEX value back.X_demo, y_demo, demo_names = load_credit_default(include_demographics=True)_, X_demo_test, _, _ = train_test_split( X_demo, y_demo, test_size=0.25, stratify=y_demo, random_state=SEED)sex = X_demo_test[:, demo_names.index("SEX")].astype(int).tolist()assert "SEX" not in feature_names # §6 and §8 need the per-feature contributions, which decide() leaves out# by default.decisions = [decide(artifact, r.tolist(), include_contributions=True) for r in X_test] # The cutoff is a policy, not a statistic: here, approve bands G01–G07.bands = artifact["bands"]cutoff = bands["edges_int"][bands["labels"].index("G08")] audit = FairnessAudit( decisions, y_test, sex, labels={1: "Male", 2: "Female"}, threshold_int=cutoff, artifact=artifact, X=X_test, protected_feature="SEX", # named, so §11 can say whether the model uses it)audit.compute_all()audit.print_summary() # The decomposition is exact because each decision's contributions sum to its# score's distance from the baseline, in integers. Summed over a group, the# identity still holds, so the per-feature gaps sum to the group gap.for code, name in ((1, "Male"), (2, "Female")): rows = [d for d, s in zip(decisions, sex) if s == code] movement = sum(2 * (d["raw_micro"] - d["baseline_micro"]) for d in rows) parts = sum(c["impact_half_micro"] for d in rows for c in d["contributions"]) print(f"{name:<7} movement {movement:>13,} contributions {parts:>13,} equal {movement == parts}") print()print("§11 counterfactual:", audit.section(11)["status"])Output#
Captured from an actual run against compileml 0.9.0 and the UCI credit panel. If this script stops working, the build fails.
captured in CIfairness_audit.py
Fairness audit============================================================== Male (n=2,983 39.8%) observed bad rate : 23.53% approval rate : 69.09% [67.41%, 70.72%] FPR / FNR : 21.39% / 38.18% near the cutoff : 12.4% within 25 points Female (n=4,517 60.2%) observed bad rate : 21.19% approval rate : 70.95% [69.61%, 72.26%] FPR / FNR : 20.51% / 39.18% near the cutoff : 13.4% within 25 points adverse impact ratio (Male / Female): 0.974 — inside the 0.80-1.25 window mean score gap decomposed exactly (residual 0); largest drivers: LIMIT_BAL 35.3% of the gap PAY_2 23.5% of the gap PAY_5 11.9% of the gap adverse-action reason divergence (total variation): 0.050 This is evidence, not a verdict. It does not certify compliance with ECOA, Regulation B, or anything else.==============================================================Male movement 721,763,812 contributions 721,763,812 equal TrueFemale movement 1,032,896,530 contributions 1,032,896,530 equal True §11 counterfactual: not_an_inputNotes#
- The audit reads decide() payloads, so it examines the decisions production made and the reason codes an applicant would be sent, not a re-scoring of the model.
- The cutoff is yours to set. Leaving threshold_int out defaults to the median score, which is a reporting convenience rather than a policy.
- SEX is named as protected_feature and is not a model input, so §11 reports not_an_input: the model cannot use it directly. Leaving protected_feature out reports not_named instead, which is not evidence either way.
- Above depth 2 attribution is no longer exact, and the gap decomposition raises instead of approximating.