October 3, 2026
AI judges need evaluation too: Chip Huyen proposes checking their scores against human ratings
Automated response evaluation needs its own checks: rising AI judge scores can diverge from expert ratings.

According to Chip Huyen's observations on Jan 16, 2025, teams behind leading AI products manually review 30–1,000 responses daily alongside automated evals.
Evals assess the quality of an AI application before it reaches users. They use language model metrics, exact checks of results, and an AI judge that evaluates another model's responses. Comparative evaluation ranks models by quality on selected tasks.
Checking the judge. Huyen recommends comparing human and automated ratings. If experts give increasingly lower scores while the AI judge gives increasingly higher ones, the judge itself needs checking. Its quality depends on the model, prompt, and task.
Checking the agent. In her Jan 7, 2025 analysis, Huyen proposes collecting tasks with lists of available tools and generating several plans for each task. Then calculate 6 metrics, including the proportion of valid plans and tool call errors. She also suggests tracking the number of steps, execution cost, and action duration, comparing the agent with another agent or a human.
Early improvements do not set the pace for the entire effort. According to Huyen's Jan 16, 2025 account, LinkedIn reached 80% of its desired AI product quality in a month. Getting past 95% took another 4 months, and the early success led the team to underestimate the difficulty of further improvements.
In the LinkedIn example, bringing the product up to the desired quality took longer than the initial improvements.