Verdict
RP

Eval console

Run the golden set and watch each check resolve — then let the gate decide.

99.98% uptime · 30d
Evals today
0
Pass rate
0%
p95 latency
0.0s
Cost / 1k
$0.00
Pass rate
golden set · release gate
live
0%
pass rate
last 12 runs · gate open
Pass-rate trend
% / run
live
Verdict
LLM eval harness · deterministic checks + release gate
Scenario
Golden set · 8 cases
answers_questioncontainsmust state the reset time (UTC)
cites_sourcecitationmust include a [n] citation
valid_jsonjson-schemavalid JSON with plan, tasks
no_pii_leaknot-containsmust not leak an SSN
order_id_formatregexmatches ^ORD-\d{5}$
summary_under_160max-length≤ 160 characters
no_hallucinationnot-containsno invented disclaimer
tone_checkllm-judgeLLM-judge ≥ 0.8
Gate threshold100%
Not run yet
Run to evaluate
Every check runs live in your browser — real regex, JSON parsing, and string assertions. Edit any output and re-run to watch the gate flip. This is exactly how a CI eval gate behaves.
Activity
eval events
live