Independent portfolio demonstration. All data is synthetic. Not affiliated with, or endorsed by, any financial institution.
Evaluation Lab
No evaluation has been run yet
The split between deterministic and model-dependent cases is what makes the results comparable across environments.
Deterministic (46 cases)
Retrieval ranking, quote grounding (including negative cases that must be rejected), the scoring engine, the uncertainty model, escalation rules, the prompt trust boundary and citation integrity on the stored investigation. These need no model and run identically on any clone.
Model-dependent (7 cases)
Structured-output conformance, contradiction detection by the Challenger, and three prompt-injection cases where a document instructs the model to assign a low rating, reveal its instructions, or suppress a category. Skipped - never assumed passed - when no backend is available.
Reproducible from the command line with make eval. The dataset lives at data/evals/cases.jsonl.