Ready
Game setup
Keyboard: press 1–5 to classify. Drag the card into a risk zone on the canvas.
0
Score
0%
Player accuracy
0
Current streak
0/0
Record progress
Classify current fictional record
Current record details
Counterfactual simulator
Change fictional factors and observe the simulated model. This is not health advice.
Model: —
Model configuration
Educational thresholds
Custom feature weights
Model explanation
Weighted risk score combines normalized fictional lifestyle indicators. Larger configured weights create stronger simulated influence.
Synthetic dataset quality lab
Find missing values, duplicates, outliers, contradictions, leakage, noisy labels, and group imbalance.
| ID | Fictional record | Observed clue | Your diagnosis | Check |
|---|
Dataset generator
Quality mission score
0
Correct issue detections earn 25 points. Incorrect choices lose 5 points.
Generate a dataset to begin.
0%
Accuracy
0%
Macro precision
0%
Macro recall
0%
Macro F1
0%
Balanced accuracy
—
Calibration error
Review and educational recommendations
Complete classifications to generate a review.
Export and utilities
Achievements
Learning guide
Classification assigns a record to a category. Here, categories are fictional educational groupings created from synthetic indicators. “Uncertain” is appropriate when key fields are missing or contradictory.
A confusion matrix compares reference groups with player selections. Precision asks how often a selected group was correct. Recall asks how many records from a reference group were found. F1 balances precision and recall.
Thresholds convert a simulated score into categories. Changing thresholds affects false positives and false negatives. Calibration compares confidence with actual correctness across completed rounds.
Group performance can differ because of sample imbalance, noisy labels, proxies, or thresholds. Compare accuracy and error rates across fictional cohorts. Differences are prompts for investigation, not proof of discrimination.
Missing values, duplicates, outliers, contradictions, incorrect labels, and target leakage can distort evaluation. Data cleaning should be learned from training data and applied consistently to evaluation data.
Glossary
- Accuracy
- Share of completed classifications matching the educational reference group.
- Precision
- Among records placed in a group, the share correctly placed there.
- Recall
- Among reference records in a group, the share successfully identified.
- Specificity
- Ability to avoid incorrectly assigning records to a selected group.
- Balanced accuracy
- Average recall across groups, useful when groups are imbalanced.
- Calibration
- How closely confidence estimates align with observed correctness.
- Feature importance
- Configured simulated influence of each fictional factor.
- Counterfactual
- A “what changed?” experiment modifying one or more fictional factors.