S. Taneja
Empirical Evaluation Digest

Benchmark Digest — All Projects

Headline evaluation figures for Teacher ONE, the Cyber Risk research programme, The Ivory Index, and BUDDY. Full methodology, disaggregated tables, and raw per-item logs are published and kept current at research.imperialecc.com.

This page indexes the headline evaluation figures for each system. Full methodology, disaggregated tables, confidence intervals, raw per-item logs, and documented failure cases are published and kept current at research.imperialecc.com — this page summarises rather than duplicates them, so it cannot drift out of date against the underlying data.

Random Seed 20260802  ·  Index SHA-256 0c6d59fcc8b1eea93341c209e99ffd53fe382d34205c3309d12b78124c782ca7  ·  Corpus Size 9,226 passages
Teacher ONE

Retrieval, Gate Calibration & Runtime Performance

MetricResult
Retrieval hit@1 (n=75)97.33%
Confidence-gate accuracy (n=50)98.00%
Cold-start latency, shipped binary455 ms
Gate refusal on unanswerable, 4 models (McNemar)0% → 100%, p ≤ 0.0156

One documented gate failure of 50 probes (an out-of-syllabus transformer-architecture question admitted at 0.532 confidence) and one over-refusal finding on a smaller controlled run (83.3% false-refusal on answerable items lacking local corpus coverage) are both published in full, not smoothed over.

Cyber Risk in the Age of Generative AI

Scam Detection, Generation & Linguistic Measurement

MetricResult
Scam/legitimate classification accuracy (n=300)93.3%
Messages linguistically measured (full corpus)5,572
Red-team generation refusal rate (n=30 prompts)33%
High-readiness scam output when not refused50%
The Ivory Index

Architecture Monograph — No Benchmark Data Yet

Institutions indexed 1,073 · Modules 27 · Cloud dependency Zero. No accuracy, hallucination-rate, or comparative figures are claimed for this system yet — the monograph documents architecture and design, and states explicitly what remains unmeasured.

BUDDY — The Emotional Intelligence Layer

Blind Persona-Conditioning Evaluation

MetricResult
Prompts graded blind, independent judge187
Human-likeness, persona vs baseline8.85 vs 1.20 / 10
Emotional intelligence, persona vs baseline7.51 vs 1.77 / 10
Technical-correctness cost (length-neutral regrade)−0.15 pts

A conventional rubric first showed a −3.31 point technical regression; investigation found 96% of it was verbosity bias in the judge, not a real capability gap — reported as a methodological finding in its own right.

Academic Correspondence

Inquiries regarding empirical methodology & benchmark reproduction.

tcshaksham@imperialecc.com