Decision field
Each point represents a fictional case. Shape indicates outcome; ring indicates a model error.
Fictional user groups
Threshold laboratory
Pass conditions
Plotly evaluation dashboard
Outcome inspection
Fairness diagnostics
| Metric | Value | Status |
|---|
Spend budget strategically
Proxy and feature audit
Unexpected deployment change
Model tournament table
| Round | Model | Accuracy | Fairness | Score |
|---|---|---|---|---|
| No completed rounds yet. | ||||
Group comparison dashboard
| Group | Samples | Accuracy | Precision | Recall | F1 | FPR | FNR | Selection | Calibration |
|---|
Coach and activity log
Fair evaluation concepts
Group fairness
Compares rates or outcomes across defined fictional groups. Passing one definition does not guarantee every definition passes.
Representation bias
Small or unbalanced groups can produce unstable estimates. More representative data can improve reliability.
Proxy variables
A seemingly neutral feature may strongly correlate with group membership and reproduce unwanted disparities.
Metric trade-offs
Accuracy, calibration, equal opportunity, and demographic parity may conflict under different base rates.
Worst-group performance
Overall accuracy can hide weak performance. Inspect the least-served group before approving deployment.
Human review
Simulation metrics support investigation, not automatic approval. Real deployments require context, governance, and ongoing review.