Eval dashboard · autonomous-dev-loop · scorecard.json
Eval dashboard
Offline evals of the LLM pipeline stages against labelled datasets (docs, ADR-0027). Last 10 live runs per suite; runs that failed on error_rate (provider outage) are left out. Updated 2026-10-04T03:12:09.640Z.
validation ✓ Pass
Last run 2026-10-04 · run 37173049724 · model groq:openai/gpt-oss-120b · 15 cases × 3 repeats
verdict_match
1
score_in_range
1
suggested_ac_count
1
valid recall
1
invalid recall
1
error_rate
0
consistency
1
latency p95
17.4 s
tokens in / out
86.1k / 8.9k
Thresholds
| Metric | Gate | Value | Status |
|---|---|---|---|
scores.verdict_match.mean | ≥ 0.8 | 1 | ✓ pass |
per_class.invalid.recall | ≥ 0.8 | 1 | ✓ pass |
error_rate | ≤ 0.05 | 0 | ✓ pass |
Trend
The trend chart appears from the second recorded run.
History
| Date | Run | Model | Dataset | verdict_match | score_in_range | suggested_ac_count | error_rate | consistency | Gate |
|---|---|---|---|---|---|---|---|---|---|
| 2026-10-04 | 37173049724 | groq:openai/gpt-oss-120b | f37c25c0 | 1 | 1 | 1 | 0 | 1 | ✓ |