Instant demo
No keys, accounts, database, or setup beyond npm install.
Open model evaluation
Compare model behavior, score it with a visible rubric, and trace every ranking back to its inputs.
Demo ranking from fixed, reproducible fixtures.
Configure one run, inspect each response, then score the evidence.
Your selected outputs and operational metrics will appear here.
Session scores update the fixed synthetic history without persistence.
| Rank | Model | Overall | Values | p95 latency | Cost / 1k | Tokens | Runs |
|---|---|---|---|---|---|---|---|
| 01 | Cedar ReasonerSynthetic Provider A | 4.48 | 100% | 1880 ms | $0.0070 | 579 | 3 |
| 02 | Baobab BalancedSynthetic Provider B | 3.97 | 67% | 1540 ms | $0.0051 | 537 | 3 |
| 03 | Marula FastSynthetic Provider C | 3.88 | 100% | 790 ms | $0.0011 | 431 | 3 |
| 04 | Karoo CompactSynthetic Provider D | 3.28 | 67% | 1030 ms | $0.0007 | 391 | 3 |
p95 uses nearest-rank. Cost per 1k divides total illustrative cost by total tokens, then multiplies by 1,000.
Clear operating boundary
The synthetic path needs no credentials. Live keys stay server-side and results remain session-only.
No keys, accounts, database, or setup beyond npm install.
The browser calls Umbono, and Umbono calls your configured provider.
Scoring, cost, p95, tokens, and ranking are tested pure functions.
Prompts and scores are not persisted by Umbono.