attribution_concentration
from compileml.fairness import attribution_concentration
§8 — how many drivers it takes to explain a decision, per group.
This section has been reformulated twice, and both changes were forced by the object rather than by taste.
It began as **curvature**, which is the right instrument for a smooth compiling function with real manifold structure, where a Hessian means something. A depth-2 ensemble is piecewise constant: second derivatives vanish almost everywhere and finite differences measure noise.
It was then **interaction share** — main-effect points versus pairwise grid points, read off the scorecard. That is exact, and on real artifacts it is useless: a depth-2 ensemble trained on real data produced *zero* main effects and forty-seven interaction grids, because no tree happened to split on a single feature. The share is then 100% for everyone, which describes the model and says nothing about any group.
What does vary by group, and answers the original question, is how
**concentrated** a decision's explanation is. A row whose score movement
comes overwhelmingly from one driver is legible and defensible; one spread
thinly across six is neither, and it is harder to write an adverse-action
notice for. Reported as the top driver's share and as an effective number
of drivers, exp(entropy), which reads directly: "this group's
decisions are explained by 3.2 features on average, that group's by 4.1."
Uses the same exact contributions as §6, so it needs no scorecard and inherits the zero-residual guarantee.
Parameters#
| Name | Type | Default | Kind |
|---|---|---|---|
| decisions | — | required | positional |
| protected | — | required | positional |
| feature_names | — | required | positional |
| labels | — | None | keyword-only |
Returns#
dict
Raises#
- ValueError
Read from the function body and the private helpers it calls, not inferred. A function that raises nothing has no section here.