Start a game to receive your first classification scenario.
You will compare the harm caused by false positives and false negatives, choose a metric, and tune the decision threshold.
1. Choose the most suitable metric
2. Identify the more costly classification error
3. Tune the operating threshold
Animated case simulation
Circles represent cases. Shape outlines distinguish actual positives from actual negatives without relying only on colour.
Confusion matrix
Precision–recall threshold trade-off
Metric reference panel
Use the formulas and decision rules below when comparing classification risks.
| Metric | Formula | Best used when |
|---|---|---|
| Precision | TP / (TP + FP) | False positives are expensive, disruptive, or limited by review capacity. |
| Recall | TP / (TP + FN) | Missing a positive case creates the greatest safety or business risk. |
| F1-score | 2PR / (P + R) | Precision and recall both matter and one balanced score is needed. |
| Specificity | TN / (TN + FP) | Correctly rejecting negative cases is the main operational requirement. |
| Accuracy | (TP + TN) / N | Classes and error costs are reasonably balanced. |
| Balanced accuracy | (Recall + Specificity) / 2 | Classes are imbalanced and both classes need equal attention. |
| PR-AUC | Area under P–R curve | The positive class is rare and ranking quality matters across thresholds. |
| ROC-AUC | Area under ROC curve | Overall discrimination is compared across many thresholds and classes are not extremely rare. |
Custom scenario builder
Create a reusable fictional classification challenge. Saved scenarios appear in future rounds.
Performance dashboard
Review metric choices, response speed, threshold quality, and category performance.
| Round | Scenario | Chosen | Correct | Threshold | Points | Time |
|---|---|---|---|---|---|---|
| No completed rounds yet. | ||||||