compileml.fairness

attribution_concentration

def attribution_concentration(decisions, protected, feature_names, *, labels=None) -> dict

from compileml.fairness import attribution_concentration

§8 — how many drivers it takes to explain a decision, per group.

This section has been reformulated twice, and both changes were forced by the object rather than by taste.

It began as **curvature**, which is the right instrument for a smooth compiling function with real manifold structure, where a Hessian means something. A depth-2 ensemble is piecewise constant: second derivatives vanish almost everywhere and finite differences measure noise.

It was then **interaction share** — main-effect points versus pairwise grid points, read off the scorecard. That is exact, and on real artifacts it is useless: a depth-2 ensemble trained on real data produced *zero* main effects and forty-seven interaction grids, because no tree happened to split on a single feature. The share is then 100% for everyone, which describes the model and says nothing about any group.

What does vary by group, and answers the original question, is how **concentrated** a decision's explanation is. A row whose score movement comes overwhelmingly from one driver is legible and defensible; one spread thinly across six is neither, and it is harder to write an adverse-action notice for. Reported as the top driver's share and as an effective number of drivers, exp(entropy), which reads directly: "this group's decisions are explained by 3.2 features on average, that group's by 4.1."

Uses the same exact contributions as §6, so it needs no scorecard and inherits the zero-residual guarantee.

Parameters#

NameTypeDefaultKind
decisions—requiredpositional
protected—requiredpositional
feature_names—requiredpositional
labels—Nonekeyword-only

Returns#

dict

Raises#

  • ValueError

Read from the function body and the private helpers it calls, not inferred. A function that raises nothing has no section here.