Reward Integrity for AI training and evaluation

Glossary

Plain definitions of the terms used on this site, in alphabetical order.

AI judge
A model that scores answers, usually against a rubric, to produce a reward or an evaluation score. Also called an LLM judge, or LLM-as-a-judge.
Configuration drift
A change in the reward that nobody decided to make: a judge model update, an edited rubric or prompt, a re-pulled container image, a floating dependency.
Evaluation integrity
A broader term for whether an evaluation measures what it claims. It also covers problems such as contamination, where test data was seen in training, and sandbagging, where a model underperforms on purpose. Reward integrity is narrower: whether the scoring pays for the right thing, in training or in evaluation.
False fail
Correct work the reward refuses. Used here instead of “false negative”, which is ambiguous outside statistics.
False pass
Wrong work the reward pays for: a wrong fix that passes the tests, or a wrong answer an AI judge accepts. Used here instead of “false positive”, which is ambiguous outside statistics.
Flaky grader
A grader that gives the same work different rewards on different runs, because of flaky tests, timing, or sampling in an AI judge.
Grader
Whatever produces the reward: tests, an answer checker, an AI judge with a rubric, or the state checks in an agent environment.
Leak surface
The answers a model can reach from inside the task, for example a reference solution in git history, in image layers, in cached files or over the network.
Reference solution
The solution the task author provides as correct. A check that the reference passes means something only if the reference passes reliably and does real work.
Reward hacking
When a model earns high reward through behavior its designers did not intend, such as special-casing the tested inputs, editing tests, reading leaked answers, or writing answers a judge over-rewards. It usually means the reward pays for something other than the intended task.
Reward integrity
Whether the reward used to train or evaluate an AI system pays for the right thing: full credit for work that does what the task asks, no credit for work that does not, and no way to move the score except by doing the work.
Reward tampering
A form of reward hacking in which the model changes the reward process itself, such as tests, scoring scripts, logs or its own reported metrics, rather than the work being scored.
RLVR
Reinforcement learning with verifiable rewards: training in which the reward comes from an automatic check, such as unit tests or an answer checker, rather than from a learned reward model.
Search pass
One run of a search for wrong work that a grader pays for, with a stated method and budget. What a search pass finds is a floor: a deeper search can find more.
Specification gaming
Behavior that satisfies the literal specification of an objective without achieving the intended outcome. Often used interchangeably with reward hacking.
Tamper surface
The grader inputs a model can change before grading: test files, fixtures, scoring scripts, logs, reported metrics.
Verifier
A common name, especially in RLVR, for a grader that checks an answer or a solution automatically. This site says “grader” to cover tests, checkers and AI judges alike.
Wrong work paid
Plain words for a false pass caused by a grader that misses a requirement the task states.

Missing a term, or think a definition is wrong? Email shane@plumblinegrader.com. Back to the definition of reward integrity.