Independent portfolio demonstration. All data is synthetic. Not affiliated with, or endorsed by, any financial institution.
AI Operations
Investigations
1
8 steps executed
Model calls
7
0 schema retries
Tokens
39,765
32,244 in / 7,521 out
Estimated cost
US$0.0000
From a price table, not billing
Failed steps
0
100.0% success rate
Replayed steps
7
Served from recordings
Each reusable capability, with the work it did and what it cost.
| Capability | Runs | Failures | Retries | Avg duration | Tokens | Est. cost |
|---|---|---|---|---|---|---|
| Planning | 1 | 0 | 0 | 7ms | 3,067 | — |
| Evidence Extraction | 1 | 0 | 0 | 1.6s | 12,916 | — |
| Risk Analysis | 1 | 0 | 0 | 4ms | 4,877 | — |
| Policy Retrieval | 1 | 0 | 0 | 15ms | 6,391 | — |
| Challenge | 1 | 0 | 0 | 3ms | 4,050 | — |
| Verification | 1 | 0 | 0 | 2ms | 3,719 | — |
| Scoring | 1 | 0 | 0 | — | 0 | — |
| Synthesis | 1 | 0 | 0 | 5ms | 4,745 | — |
Traces record what each step did - capability, tool, retrieval counts, duration, tokens and cost. They deliberately exclude model chain-of-thought, which is neither stored nor displayed.
The most honest read on whether the system is useful.
A rate near zero suggests analysts are accepting output without scrutiny. A very high rate suggests the model is not earning its place. Both are failure modes, and neither is visible without storing the AI recommendation alongside the human decision.
How often the anti-hallucination control actually fired.
Distribution of findings across the risk taxonomy.
Measured system quality, not asserted.
make eval.