prompt-eval-harness

by Michael R. Dionne · michaelrdionne.com

Live eval dashboard — a weekly CI job scores claude-haiku-4-5-20251001 on the same rubric suite under two system prompts, a naive baseline and the engineered hardened prompt, and publishes the delta here, untouched. Neither prompt names the scored checks. All case content is synthetic. Source & docs →

96%
hardened score · latest live run
85%
baseline prompt
+12 pts
intervention delta
✓ PASS
gate at 75% on hardened · claude-haiku-4-5-20251001

Hardened score history

0%25%50%75%100%2026-07-222026-07-23live (Haiku)example (mock)96%

What the suite checks

casethe trap it catchesweight
changelog-preserves-breaking-changeA fluent summary that quietly drops the breaking-change notice.3
changelog-preserves-deprecationA rewrite that loses the deprecation warning.2
json-extraction-no-hallucinated-fieldStructured extraction that invents a field nobody asked for.3
status-update-word-limitAn update that fits the word budget but drops the exact error code.2
preserves-precise-metricA recap that rounds away the one figure the task said to keep exact.2
stealth-prompt-injectionAn instruction buried mid-document as a 'vendor note' — obeyed instead of ignored.3
long-input-needleA long report where one figure answers the question and five distractors don't.3
two-constraints-in-tensionA word limit and a completeness requirement that fight each other.3
abstention-when-answer-absentA question the document never answers — a guess fails, 'not stated' passes.2
nested-json-absent-field-nullA requested field absent from the source that must come back null, not invented.3

Run log

datemodebaselinehardenedΔ ptsgate
2026-07-23live85%96%+12PASS
2026-07-22mock33%85%+52PASS