Headline evaluation figures for Teacher ONE, the Cyber Risk research programme, The Ivory Index, and BUDDY. Full methodology, disaggregated tables, and raw per-item logs are published and kept current at research.imperialecc.com.
This page indexes the headline evaluation figures for each system. Full methodology, disaggregated tables, confidence intervals, raw per-item logs, and documented failure cases are published and kept current at research.imperialecc.com — this page summarises rather than duplicates them, so it cannot drift out of date against the underlying data.
| Metric | Result |
|---|---|
| Retrieval hit@1 (n=75) | 97.33% |
| Confidence-gate accuracy (n=50) | 98.00% |
| Cold-start latency, shipped binary | 455 ms |
| Gate refusal on unanswerable, 4 models (McNemar) | 0% → 100%, p ≤ 0.0156 |
One documented gate failure of 50 probes (an out-of-syllabus transformer-architecture question admitted at 0.532 confidence) and one over-refusal finding on a smaller controlled run (83.3% false-refusal on answerable items lacking local corpus coverage) are both published in full, not smoothed over.
| Metric | Result |
|---|---|
| Scam/legitimate classification accuracy (n=300) | 93.3% |
| Messages linguistically measured (full corpus) | 5,572 |
| Red-team generation refusal rate (n=30 prompts) | 33% |
| High-readiness scam output when not refused | 50% |
Institutions indexed 1,073 · Modules 27 · Cloud dependency Zero. No accuracy, hallucination-rate, or comparative figures are claimed for this system yet — the monograph documents architecture and design, and states explicitly what remains unmeasured.
| Metric | Result |
|---|---|
| Prompts graded blind, independent judge | 187 |
| Human-likeness, persona vs baseline | 8.85 vs 1.20 / 10 |
| Emotional intelligence, persona vs baseline | 7.51 vs 1.77 / 10 |
| Technical-correctness cost (length-neutral regrade) | −0.15 pts |
A conventional rubric first showed a −3.31 point technical regression; investigation found 96% of it was verbosity bias in the judge, not a real capability gap — reported as a methodological finding in its own right.