Research /
Quantifying Label-Induced Bias in Large Language Model Self- and Cross-Evaluations
Experiments with ChatGPT, Gemini, and Claude show that authorship labels alone can materially change model judgments, supporting blind and diverse multi-model evaluation.

AI EvaluationPublished by arXiv
The study asks ChatGPT, Gemini, and Claude to judge model-written blog posts while the evaluator sees true, false, or hidden authorship labels. The content stays the same while only the identity cue changes.
Those labels shifted votes by as much as 50 percentage points and ratings by as much as 12 points. The findings support blind evaluation and a diverse panel of models when AI systems are used as judges.
Have a real process in mind?
Turn operational friction into a practical AI project.
Tell us where work is slow, repetitive, or hard to scale. We’ll help you find the smallest useful place to start.