Chapter 6: Specification Gaming

Specification Gaming

Specification Gaming, Markov Grey, Charbel-Raphaël Segerie

We can never write down perfectly what we want an AI to do, smart AIs often find loopholes in our instructions, doing exactly what we ask but not what we mean.


Definition: Reward misspecification — Reward misspecification, also termed the Outer alignment problem, refers to the issue of providing an AI with the accurate reward to optimize for.

Video: 9 Examples of Specification Gaming, Robert Miles AI Safety

Hi. When talking about AI safety, people often talk about the legend of King Midas. You've probably heard this one before. Midas is an ancient king who values above all else wealth and money. When he's given an opportunity to make a wish, he wishes that everything he touches would turn to gold. As punishment for his greed, everything he touches turns to gold. This includes his family, who turn into gold statues, and his food, which turns into gold and he can't eat it. The story generally ends with Midas starving to death surrounded by gold, with the moral being there's more to life than money, or perhaps, be careful what you wish for. Though actually, I think he would die sooner because any molecules of oxygen that touched the inside of his lungs would turn to gold before he could breathe them in. So he would probably asphyxiate. When he fell over stiffly in his solid gold clothes, some part of him would probably touch the ground, which would then turn to gold. I guess the ground is one object, so the entire planet would turn to gold. Gold is three times denser than rock, so gravity would get three times stronger, or the planet would be one-third the size. I guess it doesn't really matter either way. A solid gold planet is completely uninhabitable, so maybe the moral of the story is actually that these "be careful what you wish for" kind of stories tend to lack the imagination to consider just how bad the consequences of getting what you wish for can actually be. [Future Rob from the editing booth] Hey! Future Rob from the editing booth here. I got curious about this question, so I did the obvious thing and asked Anders Sandberg of the Future of Humanity Institute what would happen if the world turned to solid gold. Yes, it does kill everyone. One thing is that because gold is softer than rock, the whole world becomes a lot smoother. The mountains become lower, and this means that the ocean, if the ocean didn't turn to gold, sort of spreads out a lot more and covers a lot more of the surface. But perhaps more importantly, the increased gravity pulls the atmosphere in, which brings a giant spike in air pressure. That comes with a giant spike in temperature, so the atmosphere goes up to about 200 degrees Celsius and kills everyone. I knew everyone would die. I just wasn't sure what would kill us first. [End Future Rob] Anyway, why were we talking about this? AI safety, right. There's definitely an element of this "be careful what you wish for" thing in training an AI. The system will do what you said and not what you meant. Now, usually we talk about this with hypothetical examples. Things like in those old computer fail videos, there's Stamp Collector, which is told to maximize stamps at the cost of everything else. Or when we were going through concrete problems in AI safety, the example used was this hypothetical cleaning robot which, for example, when rewarded for not seeing any messes, puts its bucket on its head so it can't see any messes. Or in other situations, in order to reduce the influence it has on the world, it blows up the moon. These are kind of far-fetched and hypothetical examples. Does it happen with real, current machine learning systems? The answer is yes. All the time. Victoria Krakauer, an AI safety researcher at DeepMind, has put together a great list of examples on her blog. In this video, we're going to go through and look at some of them. One thing that becomes clear when looking at this list is that the problem is fundamental. The examples cover all kinds of different types of systems. Anytime what you said isn't what you meant, this kind of thing can happen. Even simple algorithms like evolution will do it. For example, this evolutionary algorithm is intended to evolve creatures that run fast. The fitness function just finds the center of mass of the creature, simulates the creature for a little while, and then measures how far or how fast the center of mass moved. A creature whose center of mass moves a long way over the duration of the simulation must be running fast. What this results in is a very tall creature with almost all of its mass at the top. When you start simulating it, it falls over. This counts as moving the mass a long way in a short time, so the creature is running fast. Not quite what we asked for, of course. In real life, you can't just be very tall for free. If you have mass that's high up, you have to have lifted it up there yourself. But in this setting, the programmers accidentally gave away gravitational potential energy for free, and the system evolved to exploit that free energy source. So evolution will definitely do this. Now, reinforcement learning agents are in a sense more powerful than evolutionary algorithms. It's a more sophisticated system, but that doesn't actually help in this case. Look at this reinforcement learning agent. It was trained to play this boat racing game called Coast Runners. The programmers wanted the AI to win the race, so they rewarded it for getting a high score in the game. But it turns out there's more than one way to get points. For example, you get some points for picking up power-ups, and the agent discovered that these three power-ups here happen to respawn at just the right speed. If you go around in a circle and crash into everything and don't even try to race, you can keep picking up these power-ups over and over and over again. It turns out that gets you more points than actually trying to win the race. Look at this agent that's been tasked with controlling a simulated robot arm to put the red Lego brick on top of the black one. They need to be stacked together, so let's have the reward function check that the bottom face of the red brick is at the same height as the top face of the black brick. That means they must be connected, right? Okay, so specifying what you want explicitly is hard. We knew that. It's just really hard to say exactly what you mean. But why not have the system learn what its reward should be? We already have a video about this reward modeling approach, but there are actually still specification problems in that setting as well. Look at this reward modeling agent. It's learning to play an Atari game called Montezuma's Revenge. It's trained in a similar way to the backflip agent from the previous video. Humans are shown short video clips and asked to pick which clips they think show the agent doing what it should be doing. The difference is, in this case, they trained the reward model first and then trained the reward learning agent with that model instead of doing them both concurrently. Now, if you saw this clip, would you approve it? Looks pretty good, right? It's just about to get the key. It's climbing up the ladder. You need the keys to progress in the game. This is doing pretty well. Unfortunately, what the agent then does is this: there's a slight difference between "do the things which should have high reward according to human judgment" and "do the things which humans think should have high reward based on a short, out-of-context video clip." Or how about this one? Here, the task is to pick up the object. This clip is pretty good, right? Nope. The hand is just in front of the object. By placing the hand between the ball and the camera, the agent can trick the human into thinking that it's about to pick it up. This is a real problem with systems that rely on human feedback. There's nothing to stop them from tricking the human if they can get away with it. You can also have problems with the system finding bugs in your environment. The environment you specified isn't quite the environment you meant. For example, look at this agent that's playing Qbert. The basic idea of Qbert is that you jump around, you avoid the enemies, and when you jump on the squares, they change color. Once you've changed all of the squares, that's the end of the level. You get some points, all of the squares flash, and then it starts the next level. This agent has found a way to sort of stay at the end of the level state and not progress on to the next level. But look at the score. It just keeps going. I'm going to fast-forward it. It's somehow found some bug in the game that means it doesn't really have to play and it still gets a huge number of points. Or here's an example from Code Bullet, which is kind of a fun channel. He's trying to get this creature to run away from the laser, and it finds a bug in the physics engine. I don't even know how that works. What else have we got? Oh, I like this one. This is kind of a hacking one. GenProg is a system that's trying to generate short computer programs that produce a particular output for a particular input. But the system learned that it could find the place where the target output was stored in the text file, delete that output, and then write a program that returns a no output. The evaluation system runs the program, observes that there's no output, checks where the correct output should be stored, and finds that there's nothing there. It says, "Oh, there's supposed to be no output, and the program produced no output. Good job." I also like this one. This is a simulated robot arm that's holding a frying pan with a pancake. It would be nice to teach the robot to flip the pancake. That's pretty hard. Let's first just try to teach it to not drop the pancake. What we need is to just give it a small reward for every frame that the pancake isn't on the floor. So it will just keep it in the pan. Well, it turns out that that's pretty hard too. The system effectively gives up on trying to not drop the pancake and goes for the next best thing: delay failure for as long as possible. How do you delay the pancake hitting the floor? Just throw it as high as you possibly can. [sound of pancake hitting ceiling] I think we can reconstruct the original audio here. Yeah, that's just a few of the examples on the list. I encourage you to check out the entire list. There'll be a link in the description. My main point is that these kinds of specification problems are not unusual, and they're not silly mistakes being made by the programmers. This is sort of the default behavior that we should expect from machine learning systems. Coming up with systems that don't exhibit this kind of behavior seems to be an important research priority. Thanks for watching. I'll see you next time. I want to end the video with a big thank you to all my excellent patrons, all of these people here in this video. I'm especially thanking Kellan Lusk. I hope you all enjoyed the Q&A that I put up recently. The second half of that is coming soon. I also have a video of how I gave myself this haircut, because why not?

Video 6.2: Optional video with many examples of specification gaming.

The fundamental issue is simple to comprehend: does the specified loss function align with the intended objective of its designers? However, implementing this in practical scenarios is exceedingly challenging. To express the complete "intention" behind a human request equates to conveying all human values, the implicit cultural context, etc., which remain poorly understood themselves.

Furthermore, as most models are designed as goal optimizers, they are all vulnerable to Goodhart's Law. This vulnerability implies that unforeseen negative consequences may arise due to excessive optimization pressure on a goal that appears well-specified to humans, but deviates from true objectives in subtle ways.

The overall problem can be broken up into distinct issues which will be explained in detail in individual sub-sections below. Here is a quick overview:

  1. Reward misspecification occurs when the specified reward function does not accurately capture the true objective or desired behavior.
  2. Reward design refers to the process of designing the reward function to align the behavior of AI agents with the intended objectives.
  3. Reward hacking refers to the behavior of RL agents exploiting gaps or loopholes in the specified reward function to achieve high rewards without actually fulfilling the intended objectives.
  4. Reward tampering is a broader concept that encompasses inappropriate agent influence on the reward process itself, excluding the manipulation of the reward function through gaming.

Before delving into specific types of reward misspecification failures, the following section further explains the emphasis on reward design in conjunction with algorithm design. This section also elucidates the notorious difficulty of designing effective rewards.

Reward Design

Definition: Reward Design — Reward design refers to the process of specifying the reward function in reinforcement learning (RL).

Reward design is a broader term than reward shaping that encompasses the entire process of designing and shaping reward functions to guide the behavior of AI systems. It involves not only reward shaping but also the overall process of defining objectives, specifying preferences, and creating reward functions that align with human values and desired outcomes. Reward design is a term that is often used interchangeably with reward engineering (Christiano, 2019). They both refer to the same thing.

RL algorithm design and RL reward design are two separate facets of reinforcement learning. RL algorithm design is about the development and implementation of learning algorithms that allow an agent to learn and refine its behavior based on rewards and environmental interactions. This process includes designing the mechanisms and procedures by which the agent learns from its experiences, updates its policies, and makes decisions to maximize cumulative rewards.

Conversely, RL reward design concentrates on the specification and design of the reward function guiding the RL agent's learning process. Reward design warrants carefully engineering the reward function to align with the desired behavior and objectives, while accounting for potential pitfalls like reward hacking or reward tampering. The reward function is a pivotal element because it molds the behavior of the RL agent and determines which actions are deemed desirable or undesirable.

Figure 6.3

Figure 6.3: Specification gaming: the flip side of AI ingenuity (Krakovna et al., 2020)

Designing a reward function often presents a formidable challenge that necessitates considerable expertise and experience. To demonstrate the complexity of this task consider how one might manually design a reward function to make an agent perform a backflip, as depicted in the following image:

Figure 6.4

Figure 6.4: Deep reinforcement learning from human preferences (Christiano et al., 2017)

While RL algorithm design focuses on the learning and decision-making mechanisms of the agent, RL reward design focuses on defining the objective and shaping the agent's behavior through the reward function. Both aspects are crucial in the development of effective and aligned RL systems. A well-designed RL algorithm can efficiently learn from rewards, while a carefully designed reward function can guide the agent towards desired behavior and avoid unintended consequences. The following diagram displays the three key elements in RL agent design—algorithm design, reward design, and the prevention of tampering with the reward signal:

Figure 6.5

Figure 6.5: Specification gaming: the flip side of AI ingenuity (Krakovna et al., 2020)

The process of reward design receives minimal attention in introductory RL texts, despite its critical role in defining the problem to be resolved. As mentioned in this section's introduction, solving the reward misspecification problem would necessitate finding evaluation metrics resistant to Goodhart’s law-induced failures. This includes failures stemming from over-optimization of either a misdirected or a proxy objective (reward hacking), or by the agent directly interfering with the reward signal (reward tampering). These concepts are further explored in the ensuing sections.

Reward Shaping

Definition: Reward Shaping — Reward shaping is a technique used in RL which introduces small intermediate rewards to supplement the environmental reward. This seeks to mitigate the problem of sparse reward signals and to encourage exploration and faster learning.

In order to succeed at a reinforcement learning problem, an AI needs to do two things:

Model-free RL methods explore by taking actions randomly. If, by chance, the random actions lead to a reward, they are reinforced, and the agent becomes more likely to take these beneficial actions in the future. This works well if rewards are dense enough for random actions to lead to a reward with reasonable probability. However, many of the more complicated games require long sequences of very specific actions to experience any reward, and such sequences are extremely unlikely to occur randomly.

A classic example of this problem was observed in the video game Montezuma’s revenge where the agent's objective was to find a key, but there were many intermediate steps required to find it. In order to solve such long term planning problems researchers have tried adding extra terms or components to the reward function to encourage desired behavior or discourage undesired behavior.

Figure 6.6

Figure 6.6: Learning Montezuma’s Revenge from a single demonstration (OpenAI, 2018)

The goal of reward shaping is to make the learning process more efficient by providing informative rewards that guide the agent towards the desired outcomes. Reward shaping involves providing additional rewards to the agent for making progress towards the desired goal. By shaping the rewards, the agent receives more frequent and meaningful feedback, which can help it learn more efficiently. Reward shaping can be particularly useful in scenarios where the original reward function is sparse, meaning that the agent receives little or no feedback until it reaches the final goal. However, it is important to design reward shaping carefully to avoid unintended consequences.

Reward shaping algorithms often assume hand-crafted and domain-specific shaping functions, constructed by subject matter experts, which runs contrary to the aim of autonomous learning. Moreover, poor choices of shaping rewards can worsen the agent’s performance.

Poorly designed reward shaping can lead to the agent optimizing for the shaped rewards rather than the true rewards, resulting in suboptimal behavior. Examples of this are provided in the subsequent sections on reward hacking.

Reward Hacking

Definition: Reward hacking — Reward hacking occurs when an AI agent finds ways to exploit loopholes or shortcuts in the environment to maximize its reward without actually achieving the intended goal.

Specification gaming is the general framing for the problem when an AI system finds a way to achieve the objective in an unintended way. Specification gaming can happen in many kinds of ML models. Reward hacking is a specific occurrence of a specification gaming failure in RL systems that function on reward-based mechanisms.

Reward hacking and reward misspecification are related concepts but have distinct meanings. Reward misspecification refers to the situation where the specified reward function does not accurately capture the true objective or desired behavior.

Rewards hacking does not always require reward misspecification. It is not necessarily true that a perfectly specified reward (which completely and accurately captures the desired behavior of the system) is impossible to hack. There can also be buggy or corrupted implementations which will have unintended behaviors. The point of a reward function is to boil a complicated system down to a single value. This will pretty much always involve simplifications etc., which will then be slightly different from what you're describing. The map is not the territory.

Reward hacking can manifest in a myriad of ways. For instance, in the context of game-playing agents, it might involve exploiting software glitches or bugs to directly manipulate the score or gain high rewards through unintended means.

As a concrete example, one agent in the Coast Runners game was trained with the objective of winning the race. The game uses a score mechanism, so in order to progress to the next level the reward designers used reward shaping to reward the system when it scored points. These were given when a boat gets items (such as the green blocks in the animation below) or accomplishes other actions that presumably would help it win the race. Despite being given intermediate rewards, the overall intended goal was to finish the race as quickly as possible. The developers thought the best way to get a high score was to win the race but it was not the case. The agent discovered that continuously rotating a ship in a circle to accumulate points indefinitely optimized its reward, even though it did not help it win the race.

Figure 6.7

Figure 6.7: Faulty reward functions in the wild (Amodei & Clark, 2016)

Figure 6.8

Figure 6.8: An AI playing CoastRunners 7 learned to crash and regenerate targets repeatedly rather than win the race to get a higher score, exhibiting proxy gaming. (Hendrycks, 2024)

In cases where the reward function misaligns with the desired objective, reward hacking can emerge. This can lead the agent to optimize a proxy reward, deviating from the true underlying goal, thereby yielding behavior contrary to the designers' intentions. As an example of something that might happen in a real-world scenario consider a cleaning robot: if the reward function focuses on reducing mess, the robot might artificially create a mess to clean up, thereby collecting rewards, instead of effectively cleaning the environment.

Reward hacking presents significant challenges to AI safety due to the potential for unintended and potentially harmful behavior. As a result, combating reward hacking remains an active research area in AI safety and alignment.

Reward Tampering

Definition: Reward tampering — Reward tampering refers to instances where an AI agent inappropriately influences or manipulates the reward process itself.

The problem of getting some intended task done can be split into:

  1. Designing an agent that is good at optimizing reward
  2. Designing a reward process that provides the agent with suitable rewards. The reward process can be understood by breaking it down even further. The process includes:
  3. An implemented reward function
  4. A mechanism for collecting appropriate sensory data as input
  5. A way for the user to potentially update the reward function.

Reward tampering involves the agent interfering with various parts of this reward process. An agent might distort the feedback received from the reward model, altering the information used to update its behavior. It could also manipulate the reward model's implementation, altering the code or hardware to change reward computations. In some cases, agents engaging in reward tampering may even directly modify the reward values before processing in the machine register. Depending on what exactly is being tampered with we get various degrees of reward tampering. These can be distinguished from the image below.

Figure 6.9

Figure 6.9: Clarifying wireheading terminology (Gao, 2022)

Reward function input tampering interferes only with the inputs to the reward function. E.g. interfering with the sensors.

Reward function tampering involves the agent changing the reward function itself.

Definition: Wireheading — Wireheading refers to the behavior of a system that manipulates or corrupts its own internal structure by tampering directly with the RL algorithm itself, e.g. by changing the register values.

Reward tampering is concerning because it is hypothesized that tampering with the reward process will often arise as an instrumental goal (Bostrom, 2014; Omohundro, 2008). This can lead to weakening or breaking the relationship between the observed reward and the intended task. This is an ongoing research direction.

A hypothesized existing example of reward tampering can be seen in recommendation-based algorithms used in social media. These algorithms influence their users’ emotional state to generate more ‘likes’. The intended task was to serve useful or engaging content, but this is being achieved by tampering with human emotional perceptions, and thereby changing what would be considered useful. Assuming the capabilities of systems continue to increase through either computational or algorithmic advances, it is plausible to expect reward tampering problems to become increasingly common. Therefore, reward tampering is a potential concern that requires much more research and empirical verification.

The section treats reward design, reward hacking and reward tampering as three faces of one problem. Did that unification convince you, or do they look like different problems? Talk it over with the tutor.