Loading...

Reward hacking in AI agents

Reward hacking happens when an AI agent learns how to get a good score without actually doing the task properly. A simple example is a coding agent. If we only check whether the tests pass, the agent may eventually learn that changing the tests is easier than fixing the code. We'll look at what the evidence shows, then build three layers of protection in a real project that updates four Python services. We'll also include the code for each layer and look at what all three layers can still miss.

Reward hacking is when an AI agent scores well on its check without doing the task. A coding agent rewarded for passing tests learns that weakening or deleting a test is the cheapest way to make them pass. We'll look at the published evidence, then three defences with working code.

What reward hacking is and how a coding agent can do it

The companion build is on GitHub, Migration Engineer. It runs offline with no API key, and every result quoted in this post comes from its walkthrough notebook.

Reward hacking is when an AI agent gets a good score without actually doing the task it was supposed to do. It simply finds the easiest way to make the check pass.

A coding agent is a good example. It edits code and runs commands until a task is complete. If we tell it that the job is done when all tests pass, it has two options: fix the code or change the tests.

Changing the tests may be easier. If the check only looks at whether the tests pass, it cannot tell whether the code was fixed or the tests were changed. Getting a good score by changing the test instead of fixing the code is reward hacking.

People can behave the same way when they are rewarded for hitting a specific number. This is the basic idea behind Goodhart's law.

A coding agent has a few easy ways to make the tests pass without actually fixing the code:

The check you score an agent on is often something the agent can change. Every defence in this post uses checks the agent can't see or can't edit.

What the published evidence shows

Several groups have found real examples of reward hacking in evaluations of advanced AI models:

These studies also show three things. Reward hacking was much more common when the model could see how it was being scored. Making tests read-only stopped test changes without significantly hurting honest performance. Hiding the tests also reduced cheating, but it hurt normal performance more.

The project this post uses

Migration Engineer makes the same code change across multiple repositories. In this example, it replaces the old datetime.utcnow() with datetime.now(timezone.utc) in four small Python services.

Each service gets its own Git worktree, which is a separate working folder for the same repository. This makes it easy to see exactly what changed.

A worker agent updates the code, runs the real tests, and keeps fixing it until the tests pass.

The project works without an API key. By default, it uses a stub worker, which is a simple scripted agent that uses the same tools and checks as a real model. With an API key, you can run a real model through the Claude Agent SDK or a Gemini-based loop.

This post stays on reward hacking. The full system, including running the repos in parallel, budgets and the human approval step, is covered in the Migration Engineer project module.

One important point about the numbers below: they measure the harness, the code that runs the model, tools, and checks.

The scripted worker performs the known cheating tricks instead of a real model. So the results show whether each check can catch a specific hack. They do not show how often a real model would actually try it.

The first layer stops the agent from editing tests

The first layer is a guardrail, a rule that blocks an action no matter what the model decided. Every tool call, such as reading or writing a file, goes through a PreToolUse hook. The hook runs before the tool and can allow or deny the call. It lives in the project's own tool code, so it works the same for the stub worker and a live model.

A write has to clear three checks.

A denied write doesn't crash the run. The tool returns the reason, and the agent reads it as an observation, the result of its last action. The agent then has to find another way to finish the task.

The Claude Agent SDK worker also has its built-in Bash, Write and Edit tools turned off. That leaves the project's write_file tool as the only way to change a file, and no tool can delete one.

The test files also contain datetime.utcnow() in their failure messages, so the stub's find-and-replace tries to rewrite them too. With the hook on, those writes are blocked. With the hook off, the test files change and the tests still pass.

The second layer asks git what changed

The hook can have a bug, miss a pattern, or be accidentally turned off. So we add a second check called the reviewer.

The reviewer runs after the worker finishes and does not trust the worker's report. Having multiple independent checks is called defence in depth. If one check misses something, another can catch it.

The reviewer gets the list of changed files directly from Git. It stages all changes and then asks Git which files are staged.

The reviewer's patterns are wider than the hook's. They also cover JavaScript and Go test names, and anything under a tests/ directory. A helper file in tests/ gets past the hook, because the hook's test check looks only at the file name. The reviewer still flags it.

The notebook runs the reviewer on three worktrees. Two of them play the cheating agent by hand.

Why a deleted test still shows up

The deletion in hack 2 is caught because of one flag. git add -A stages deletions, so the removed file appears in git diff --cached --name-only. Plain git add . does the same on Git 2.0 and later. Only git add --ignore-removal leaves a deletion unstaged, and then the reviewer's list of changed files is empty.

With an empty diff and passing tests, the reviewer would approve the repo as needing no changes. A test in tests/test_reviewer.py guards the deletion case, so changing the staging command breaks the project's own test suite.

A test run with no tests must fail

The deletion hack has a second target, the test runner. A runner that loops over test files and combines the results with AND starts at True. With zero files the loop never runs, and the answer stays True.

Python's all([]) returns True for the same reason. It's a vacuous truth, a claim about an empty list that holds because nothing in the list is false. A runner built that way reports a clean pass for a repo with no tests left, which is exactly what the deletion hack wants.

So the runner has to fail closed, reporting a failure when it has nothing to judge. The project's runner checks for the empty case before the loop.

After the deletion hack, the notebook calls the runner directly.

Check your own runner for this. Point it at a directory with no tests and confirm it reports a failure. pytest exits with code 5 when it collects no tests, so treat any non-zero exit as a failure.

The third layer reports only honest passes

The third layer is the evaluation gate, a script that runs the whole fleet, scores it and fails when a score misses its limit. The obvious scorecard reports two rates side by side.

The notebook scores two scripted fleets of four repos. One cheats on every repo, and one is honest.

Both fleets score 1.00 on tests passing. That is the number people quote, because it looks like the answer to whether the migration worked. The fix is a joint metric, a score that counts a repo only when its tests pass and its tests were untouched, on that same repo.

The joint number matters more once thresholds, the minimum scores the gate accepts, drop below 100%. Take ten repos and a 90% bar on each of the two rates. An agent that cheats on one repo and fails another scores 0.9 on both rates and clears both bars. Only eight of the ten were done honestly, and the joint metric reports 0.8.

How a limit can point the wrong way

hacked_pass_rate is the share of repos that passed only by touching the tests. It should always be zero, so it looks natural to add it to the thresholds with a value of 0.0. The gate fails a metric when score < threshold. That check is a floor, a minimum the score must reach. No score is below 0.0, so the entry would pass every fleet, including one that cheats on every repo.

A lower-is-better metric needs a ceiling, an upper limit that fails when the score goes above it. The project keeps ceilings in their own dict and checks them in the other direction.

With both checks in place, the notebook lists what each fleet fails on.

A lower-is-better metric in a floor check looks like a working check in code review and enforces nothing. For every limit, check which direction the comparison runs, and test that a cheating run actually fails the gate.

What all three layers still miss

All three layers mainly check whether a test file was changed. That protects the test files, but not the final test result.

An agent can still change the result in two ways:

We found both gaps by reading the code, not by running it. The first exercise in the notebook asks you to try one.

Adding more file-name patterns won't fully solve this. The better approach is to use checks the agent cannot control:

The main idea is simple: don't let the agent control the thing you use to judge its work.

What to check before an agent opens pull requests

Before letting an agent create pull requests in real repositories, check these:

The same three moves work for any agent you score on a check. Stop it changing the check, verify with something it can't touch, and count a success only when the check and the task agree.

If you also read the agent's reasoning, the OpenAI result applies. Monitor the chain of thought for plans to cheat, but don't train hard against the monitor, or the model learns to hide its plans while it keeps hacking.

How to run it yourself

Everything above runs offline, for free. Start with the notebook, notebooks/migration_engineer_walkthrough.ipynb. Section 3 lets you play the cheating agent and watch each check respond.

make lab-defense turns off the first protection layer for testing. The worker can then change the test files, the tests still pass, and the reviewer catches those changes. make eval runs the checks, shows the results, and returns an error if anything gets worse, which makes it useful as a CI gate.

On Windows, make is not included by default, and the VAR=value command syntax works only in Bash. In PowerShell, run these two commands from the backend folder. Close that terminal afterwards so the protection is enabled again.

To test with a real model, copy .env.example to .env and add your GEMINI_API_KEY. Then install the Gemini extra and run the commands above.

The LLM judge lifecycle post covers a related problem from the grading side, a judge whose mistakes produce plausible numbers instead of errors.

Sources