Examples · Validation & governance

FairnessAudit

Audit the compiled decisions by a protected attribute the model was never given, and decompose the group score gap by feature exactly.

Code#

fairness_audit.py
from _setup import SEED, X_test, artifact, feature_names, y_test
from sklearn.model_selection import train_test_split
 
from compileml.datasets import load_credit_default
from compileml.fairness import FairnessAudit
from compileml.runtime import decide
 
# The artifact was compiled without demographics. Load them separately and
# repeat the same split, so each holdout row gets its own SEX value back.
X_demo, y_demo, demo_names = load_credit_default(include_demographics=True)
_, X_demo_test, _, _ = train_test_split(
X_demo, y_demo, test_size=0.25, stratify=y_demo, random_state=SEED
)
sex = X_demo_test[:, demo_names.index("SEX")].astype(int).tolist()
assert "SEX" not in feature_names
 
# §6 and §8 need the per-feature contributions, which decide() leaves out
# by default.
decisions = [decide(artifact, r.tolist(), include_contributions=True) for r in X_test]
 
# The cutoff is a policy, not a statistic: here, approve bands G01–G07.
bands = artifact["bands"]
cutoff = bands["edges_int"][bands["labels"].index("G08")]
 
audit = FairnessAudit(
decisions,
y_test,
sex,
labels={1: "Male", 2: "Female"},
threshold_int=cutoff,
artifact=artifact,
X=X_test,
protected_feature="SEX", # named, so §11 can say whether the model uses it
)
audit.compute_all()
audit.print_summary()
 
# The decomposition is exact because each decision's contributions sum to its
# score's distance from the baseline, in integers. Summed over a group, the
# identity still holds, so the per-feature gaps sum to the group gap.
for code, name in ((1, "Male"), (2, "Female")):
rows = [d for d, s in zip(decisions, sex) if s == code]
movement = sum(2 * (d["raw_micro"] - d["baseline_micro"]) for d in rows)
parts = sum(c["impact_half_micro"] for d in rows for c in d["contributions"])
print(f"{name:<7} movement {movement:>13,} contributions {parts:>13,} equal {movement == parts}")
 
print()
print("§11 counterfactual:", audit.section(11)["status"])

Output#

Captured from an actual run against compileml 0.9.0 and the UCI credit panel. If this script stops working, the build fails.

captured in CIfairness_audit.py
Fairness audit
==============================================================
 
Male (n=2,983 39.8%)
observed bad rate : 23.53%
approval rate : 69.09% [67.41%, 70.72%]
FPR / FNR : 21.39% / 38.18%
near the cutoff : 12.4% within 25 points
 
Female (n=4,517 60.2%)
observed bad rate : 21.19%
approval rate : 70.95% [69.61%, 72.26%]
FPR / FNR : 20.51% / 39.18%
near the cutoff : 13.4% within 25 points
 
adverse impact ratio (Male / Female): 0.974 — inside the 0.80-1.25 window
 
mean score gap decomposed exactly (residual 0); largest drivers:
LIMIT_BAL 35.3% of the gap
PAY_2 23.5% of the gap
PAY_5 11.9% of the gap
 
adverse-action reason divergence (total variation): 0.050
 
This is evidence, not a verdict. It does not certify compliance
with ECOA, Regulation B, or anything else.
==============================================================
Male movement 721,763,812 contributions 721,763,812 equal True
Female movement 1,032,896,530 contributions 1,032,896,530 equal True
 
§11 counterfactual: not_an_input

Notes#

  • The audit reads decide() payloads, so it examines the decisions production made and the reason codes an applicant would be sent, not a re-scoring of the model.
  • The cutoff is yours to set. Leaving threshold_int out defaults to the median score, which is a reporting convenience rather than a policy.
  • SEX is named as protected_feature and is not a model input, so §11 reports not_an_input: the model cannot use it directly. Leaving protected_feature out reports not_named instead, which is not evidence either way.
  • Above depth 2 attribution is no longer exact, and the gap decomposition raises instead of approximating.

API reference: FairnessAudit →