Interactive demo
Dashboard
Run the golden dataset through a provider, compare two models side-by-side, and watch quality trend over time. This hosted demo is mock-only — deterministic, free, no API key.
Run an evaluation
No results yet
Pick a provider and metrics above, then Run evaluation. Each case streams in live with a pass/fail verdict and a quality-gate readout.
This is a fresh session — nothing is pre-loaded. Your run, your results.
Case results
Bold metric = gates this case's pass/fail. Others are informational — a case can PASS with a low non-gating score (e.g. a summary shouldn't fail on exact string match).
| Case | Result | Tags | Latency | Metrics |
|---|
Compare providers
Run the same dataset through two providers side-by-side and see exactly where they diverge.
Quality trend
Drag any two runs from the table below into the slots to diff them.
| # | When | Provider / model | Pass | Avg | Cost |
|---|
Peek at the golden dataset
Each case declares which metric(s) gate its pass/fail. That's why a verbose-but-correct summary can still PASS — it's gated on the judge, not exact string match.