When AI scores the same student work, the models disagree, and that gap changes which students look strongest. This tool measures the instability and reports a defensible range instead of one number that depends on the model you happened to use.
Each model scores the selected essay from 1 to 6. The shaded band is the range across models. The diamond is a blinded human reference rating.
Rank all six essays by whichever model you trust, then compare each to its rank under the human reference. The same essays move up and down depending only on the scorer.
| Rank | Essay | Score | vs human | Certainty |
|---|
The instability is the finding, not a bug. These methods turn it into measurement you can defend.
Instead of one AI score, report an interval reflecting what the models can and cannot agree on, so a ranking is asserted only when the data support it.
A blinded sample of human ratings anchors the AI scores and lets you estimate and correct model-specific bias.
Which model you use shifts the answer. Measure that sensitivity directly so findings do not quietly depend on a vendor choice.
Method draws on the Principal Investigator's research on multi-model measurement instability (NBER Working Paper 35110) and platform-selection effects in AI measurement (arXiv 2605.21743).