Model validation
Aggregate numbers from the held-out test set and a same-session repeat-scan reliability check, so you can judge overall trustworthiness before trusting any one prediction. See the project README and paper for full methodology.
Diagnosis (CN / MCI / AD), test set
Balanced accuracy: 0.631
| Predicted CN | Predicted MCI | Predicted AD | |
|---|---|---|---|
| True CN | 162 | 23 | 19 |
| True MCI | 91 | 61 | 68 |
| True AD | 5 | 27 | 147 |
MCI → AD conversion risk, test set
Trained on a small cohort (170 subjects) — treat as exploratory, not diagnosis-grade.
Balanced accuracy: 0.610
| Predicted stable | Predicted converter | |
|---|---|---|
| True stable | 94 | 25 |
| True converter | 33 | 25 |
Test-retest reliability
Same person, same session, two MRI acquisitions ~8 minutes apart — does the model agree with itself?
Agreement rate: 0.870 across 376 repeat-scan pairs.
Interpretability
Grad-CAM attention checked against a hippocampus mask — the region AD is known to affect first. Higher overlap than the mask's share of total brain volume suggests the model is using anatomically relevant signal, not just any texture correlated with the label.
Sample scan: mean 1.63% across 59 test-set scans, vs. 0.54% chance level (the mask's share of the cropped volume) -- roughly 3x chance, though modest in absolute terms