AI Research & Evaluation Tools

AI Score Stability Lab

When AI scores the same student work, the models disagree, and that gap changes which students look strongest. This tool measures the instability and reports a defensible range instead of one number that depends on the model you happened to use.

Software demonstration. Prototype built by Dr. Michelle Yin. Essays, model scores, and human references below are illustrative sample data chosen to show the method clearly.

Score one essay, four ways

Each model scores the selected essay from 1 to 6. The shaded band is the range across models. The diamond is a blinded human reference rating.

Model AModel BModel CModel DHuman

The ranking flips with the model

Rank all six essays by whichever model you trust, then compare each to its rank under the human reference. The same essays move up and down depending only on the scorer.

RankEssayScorevs humanCertainty

What this measures, and what to do about it

The instability is the finding, not a bug. These methods turn it into measurement you can defend.

Report bounds, not a point

Instead of one AI score, report an interval reflecting what the models can and cannot agree on, so a ranking is asserted only when the data support it.

Calibrate to human reference

A blinded sample of human ratings anchors the AI scores and lets you estimate and correct model-specific bias.

Account for model selection

Which model you use shifts the answer. Measure that sensitivity directly so findings do not quietly depend on a vendor choice.

Method draws on the Principal Investigator's research on multi-model measurement instability (NBER Working Paper 35110) and platform-selection effects in AI measurement (arXiv 2605.21743).

More EduWork.AI tools All tools →