Eval dashboard · autonomous-dev-loop · scorecard.json

Eval dashboard

Offline evals of the LLM pipeline stages against labelled datasets (docs, ADR-0027). Last 10 live runs per suite; runs that failed on error_rate (provider outage) are left out. Updated 2026-10-04T03:12:09.640Z.

validation ✓ Pass

Last run 2026-10-04 · run 37173049724 · model groq:openai/gpt-oss-120b · 15 cases × 3 repeats

verdict_match
1
score_in_range
1
suggested_ac_count
1
valid recall
1
invalid recall
1
error_rate
0
consistency
1
latency p95
17.4 s
tokens in / out
86.1k / 8.9k

Thresholds

MetricGateValueStatus
scores.verdict_match.mean≥ 0.81✓ pass
per_class.invalid.recall≥ 0.81✓ pass
error_rate≤ 0.050✓ pass

Trend

The trend chart appears from the second recorded run.

History

DateRunModelDatasetverdict_matchscore_in_rangesuggested_ac_counterror_rateconsistencyGate
2026-10-0437173049724groq:openai/gpt-oss-120bf37c25c011101✓