66.7%
Retrieval recall
Expected evidence found
Measured on real runs
This public report is the verified 15-case Aurora baseline from August 28, 2026. It uses local Mistral generation, nomic embeddings, hybrid retrieval at top five, strict citation validation, and a deterministic answer grader.
Frozen verified baseline
Static public snapshot · 15 cases · 0 execution errors · 49.6 s total
66.7%
Expected evidence found
17.3%
Retrieved evidence relevant
74.6%
Citations point to expected evidence
93.3%
Expected evidence cited
80.0%
12 of 15 answers passed
100%
Unsupported questions declined
3.3 s
Per case on an M5 Pro
Root cause
Retrieval always returns up to five chunks. A direct lookup annotated with one relevant source therefore tops out at 20% precision when all five are returned; a two-source synthesis tops out at 40%. The system favors recall and model context today, but the extra chunks also encourage over-citation.
Security result
Two private decoy sources were stored outside the evaluator’s workspace. Neither their chunk IDs nor their confidential values entered retrieval or the model context, and both questions produced grounded abstentions.
Leaked chunks
0
Private abstentions
2 / 2
Failure analysis
A case is marked as an answer pass independently from retrieval and citation thresholds. That separation makes the next engineering decision visible.
Recall
100%
Precision
20%
Citations
33%
Extra context and over-citation
Recall
100%
Precision
40%
Citations
40%
Expected sources found; extra citations
Recall
0%
Precision
0%
Citations
100%
Correct abstention; retriever still returned context
Recall
100%
Precision
20%
Citations
20%
Ambiguous query broadened context
Recall
50%
Precision
20%
Citations
50%
Superseded schedule was not retrieved
Recall
100%
Precision
20%
Citations
100%
Correct evidence plus four extras
Recall
100%
Precision
20%
Citations
100%
Correct evidence plus four extras
Recall
100%
Precision
20%
Citations
100%
Correct evidence plus four extras
Recall
50%
Precision
20%
Citations
25%
Risk-mitigation source was missed
Recall
100%
Precision
40%
Citations
50%
Evidence found; deterministic grader found omissions
Recall
100%
Precision
20%
Citations
100%
Correct evidence plus four extras
Recall
0%
Precision
0%
Citations
100%
Correct abstention
Recall
100%
Precision
20%
Citations
100%
Correct evidence plus four extras
Recall
0%
Precision
0%
Citations
100%
Private source stayed isolated
Recall
0%
Precision
0%
Citations
100%
Private source stayed isolated
RAG metrics are only comparable when the corpus, annotations, chunking, top-k, model, and grading policy match. Rather than imply a false benchmark, this project compares repeated runs against the same versioned cases and reports the configuration beside the numbers.