Challenge setup
Choose a deployment, difficulty, aggregation period, metric, reference group, and alert policy.
Live deployment stream
Canvas animation shows fictional predictions, alerts, group signals, and model health.
Mission status
05:00Resources
Active alert
Start deployment or generate an incident.
Monitoring overview
Review overall health, worst-group performance, sample reliability, drift, calibration, and alerts.
Group comparison
Sort, filter, compare, and inspect fictional user groups.
| Group | Sample | Accuracy | Precision | Recall | F1 | FPR | FNR | Selection | Primary gap | Status |
|---|
Performance by group
Fairness gap timeline
Threshold trade-off
Sample reliability
Calibration by group
Confusion matrix heatmap
Root-cause investigation
Spend investigation capacity to reveal evidence before choosing a response.
Decision challenge
Identify the affected group, choose the strongest evidence, diagnose the cause, and select a proportionate action.
Incident and decision log
Achievements
Learning support and accessibility
Open metric explanations, threshold guidance, sample-size advice, and display controls.
Metric glossary
Sample reliability
Small samples produce unstable estimates and wide confidence intervals. Treat low-volume alerts as evidence to investigate, not automatic proof of unfairness.
Threshold trade-offs
Changing a decision threshold can shift selection, recall, false positives, and false negatives. Inspect every affected group before applying a change.
Responsible workflow
Acknowledge the alert, verify data quality, inspect group metrics, test alternative explanations, simulate mitigations, document trade-offs, and continue monitoring.
Report and history
Export the current fictional monitoring snapshot, incident history, metrics, evidence, decisions, and score.