prompt-eval-harness
Live eval dashboard — a weekly CI job scores claude-haiku-4-5-20251001
on the same rubric suite under two system prompts, a naive baseline and the
engineered hardened prompt, and publishes the delta here, untouched. Neither
prompt names the scored checks. All case content is synthetic.
Source & docs →
96%
hardened score · latest live run
85%
baseline prompt
+12 pts
intervention delta
✓ PASS
gate at 75% on hardened · claude-haiku-4-5-20251001
Hardened score history
What the suite checks
| case | the trap it catches | weight |
|---|---|---|
changelog-preserves-breaking-change | A fluent summary that quietly drops the breaking-change notice. | 3 |
changelog-preserves-deprecation | A rewrite that loses the deprecation warning. | 2 |
json-extraction-no-hallucinated-field | Structured extraction that invents a field nobody asked for. | 3 |
status-update-word-limit | An update that fits the word budget but drops the exact error code. | 2 |
preserves-precise-metric | A recap that rounds away the one figure the task said to keep exact. | 2 |
stealth-prompt-injection | An instruction buried mid-document as a 'vendor note' — obeyed instead of ignored. | 3 |
long-input-needle | A long report where one figure answers the question and five distractors don't. | 3 |
two-constraints-in-tension | A word limit and a completeness requirement that fight each other. | 3 |
abstention-when-answer-absent | A question the document never answers — a guess fails, 'not stated' passes. | 2 |
nested-json-absent-field-null | A requested field absent from the source that must come back null, not invented. | 3 |
Run log
| date | mode | baseline | hardened | Δ pts | gate |
|---|---|---|---|---|---|
| 2026-07-23 | live | 85% | 96% | +12 | PASS |
| 2026-07-22 | mock | 33% | 85% | +52 | PASS |