MarkAlign
Anything can output a score. The hard problem is telling a grader that works from one that merely looks like it works.
Creator · Evaluation design, grader profiling, calibration policy, and the eval harness
+2 more
0.664
QWK — 92% of the human-vs-human ceiling
0.369 → 0.576
calibrated QWK on a held-out second task
3 / 3
calibrate-or-not calls the diagnosis got right
Problem
Putting an AI grader in front of student work fails on a question a score cannot answer: did it match the teacher for the right reason, or did it get lucky? A bare agreement number hides the difference between a grader with a fixable systematic bias and one that is simply noisy.
Why it matters
Read against 1.0, every grader looks broken. Two trained human raters on this ASAP set agree at QWK 0.72, so the ceiling — not perfection — is the bar. Separating score-agreement from reasoning-agreement, and systematic bias from random noise, is the judgment the system is built to make legible.
Architecture
8 stages · select to inspect
Product surface



Technical challenges
Ordinal instability
LLM judges are unstable asked to rate 0–3 but steady on binary decisions. Replacing the rating with a learned yes/no ladder, each rung pinned to a verbatim quote, moved QWK from 0.585 to 0.656.
A ladder with no floor
The checklist made the worst misses worse — within-1 agreement fell to 41% because weak essays took 0s a real teacher never gives. Counting the teacher's actual score distribution over the calibration set and stating it as a constraint fixed it without loosening the ladder.
Knowing when not to calibrate
The same scale transform that recovered a harsh, compressed grader (0.40 → 0.47) actively hurt one that already had healthy spread (0.585 → 0.478). Calibration had to become conditional on the diagnosis rather than a default step.
Tradeoffs
Checklist scoring over direct rating
Trades interpretive freedom for scale stability, and makes every point traceable to a quoted span.
Diagnosis gates calibration
A blanket calibration step would have degraded the strongest configuration; the rule only fires on a monotonic-but-shifted signature.
Single held-out split, stated as such
120 held-out essays on one split is a strong single run, not a cross-validated mean, and the write-up says so rather than implying more.
Experiments
- 01Compared graders and prompt formats against the same held-out teacher marks, isolating the model swap from the prompt-format change.
- 02Derived a per-trait score-distribution floor from the 25-essay calibration set and measured every metric moving together.
- 03Re-ran the identical pipeline on ASAP set 1 — different genre, scale, and rater pair — and predicted the calibrate/don't-calibrate call before running it.
Results
QWK 0.664 against the teacher on ASAP set 7, 92% of the 0.72 human-vs-human ceiling, learning the standard from 25 essays with no fine-tuning.
On an unseen task type, diagnosis predicted the calibration case in advance: QWK 0.369 to 0.576, mean error roughly halved.
A deterministic mock mode with a deliberate length bias, so the harness can be run end to end with no API key and still catch a real systematic error.
Lessons learned
- Counting is code's job. Noticing a distribution across 25 essays is exactly what an LLM quietly fails at.
- A calibration step is a claim about the error, not a free improvement — applied blindly it degrades a healthy grader.
- Reporting against the human ceiling rather than 1.0 changes which results look like progress.
Future work
- Establish cross-task teacher transfer, which needs the same identified grader on two assignments.
- Promote the reasoning taxonomy from method to reported result with a larger adjudicated sample.
- Cross-validate the headline number instead of relying on one held-out split.