Examples · Tuning

sweep_whitebox

Reading the retention curve: how much teacher Gini survives at each whitebox size.

Code#

sweep_whitebox.py
from _setup import X_train, teacher_scores, y_train
 
from compileml.reference import fit_reference
from compileml.tune import sweep_whitebox
 
# Retention against a teacher is one-sided: it cannot report that a handful
# of logistic coefficients would have scored higher. The floor supplies the
# other side of the comparison.
reference = fit_reference(X_train, y_train)
 
rows = sweep_whitebox(
X_train,
teacher_scores,
y_train,
trees_grid=(10, 30),
depth_grid=(1, 2),
reference=reference,
)
 
header = f"{'trees':>6}{'depth':>7}{'gini':>9}{'vs teacher':>12}{'vs floor':>10}{'exact':>7}"
print(header)
for row in rows:
print(
f"{row['n_estimators']:>6}{row['max_depth']:>7}{row['gini']:>9.4f}"
f"{row['gini_retention_pct']:>11.1f}%{row['gini_vs_reference_pct']:>9.1f}%"
f"{str(row['exact_attribution']):>7}"
)
 
print()
print("Spend on trees. Be stingy with depth: above 2 the exactness goes.")

Output#

Captured from an actual run against compileml 0.9.0 and the UCI credit panel. If this script stops working, the build fails.

captured in CIsweep_whitebox.py
trees depth gini vs teacher vs floor exact
10 1 0.4726 69.5% 85.9% True
30 1 0.5467 80.4% 99.4% True
10 2 0.5378 79.0% 97.8% True
30 2 0.5620 82.6% 102.2% True
 
Spend on trees. Be stingy with depth: above 2 the exactness goes.

Notes#

  • The curve usually flattens early. Pick the smallest configuration on the flat part — smaller artifacts explain better and export smaller.

API reference: sweep_whitebox →