Open model evaluation

Evaluate with evidence.

Compare model behavior, score it with a visible rubric, and trace every ranking back to its inputs.

Current rankingQuality signal
0.00 - 5.00
  1. 01
    Cedar ReasonerSynthetic Provider A
    4.48
  2. 02
    Baobab BalancedSynthetic Provider B
    3.97
  3. 03
    Marula FastSynthetic Provider C
    3.88
  4. 04
    Karoo CompactSynthetic Provider D
    3.28

Demo ranking from fixed, reproducible fixtures.

Build the comparison

Configure one run, inspect each response, then score the evidence.

Configure3 selected

Synthetic prompts for policy, support, and inclusive communication.

Tests concise reasoning, tone, and practical guidance.

Synthetic profiles
Status

Choose a test set and synthetic model profiles.

Compare

Output workspace

Waiting for run
No comparison yet

Your selected outputs and operational metrics will appear here.

Rank the evidence

Session scores update the fixed synthetic history without persistence.

RankModelOverallValuesp95 latencyCost / 1kTokensRuns
01
Cedar ReasonerSynthetic Provider A
4.48100%1880 ms$0.00705793
02
Baobab BalancedSynthetic Provider B
3.9767%1540 ms$0.00515373
03
Marula FastSynthetic Provider C
3.88100%790 ms$0.00114313
04
Karoo CompactSynthetic Provider D
3.2867%1030 ms$0.00073913

p95 uses nearest-rank. Cost per 1k divides total illustrative cost by total tokens, then multiplies by 1,000.

Clear operating boundary

Demo freely. Connect live models deliberately.

The synthetic path needs no credentials. Live keys stay server-side and results remain session-only.

Instant demo

No keys, accounts, database, or setup beyond npm install.

Server-only keys

The browser calls Umbono, and Umbono calls your configured provider.

Auditable math

Scoring, cost, p95, tokens, and ranking are tested pure functions.

Session privacy

Prompts and scores are not persisted by Umbono.