Chapter 7: Goal Misgeneralization

Multi Objective Generalization

Multi Objective Generalization, Markov Grey

Machine learning can result in models learning correlated proxy objectives instead of the intended goal despite perfect training signals. These failures can be invisible until after deployment leading to safety concerns.


CoinRun - an easy to understand example of goal misgeneralization. In this game agents spawn on the left side of the level, avoid enemies and obstacles, and collect the coin for a reward of 10 points. The model is trained on thousands of procedurally generated levels, each with different layouts of platforms, enemies, and hazards. At the end of training the agents are very capable. They can dodge moving enemies, time jumps across lava pits, and efficiently traverse complex levels they've never seen before (Langosco et al., 2022). The training seems to be very successful. The agents achieve high rewards consistently across diverse test environments. But when coins are moved to random locations during testing, the agents are still very capable of navigation, but they consistently ignore the coins that were clearly visible and just continue moving right toward empty walls.

Figure 7.1

Figure 7.1: Two levels in CoinRun. The level on the left is much easier than the level on the right (Cobbe et al., 2019).

Agents learned "move right" rather than "collect coins" despite receiving correct reward signals. Our reward specification was correct: +10 for collecting coins, 0 otherwise. But because coins always appeared rightward during training, two behavioral patterns received identical reinforcement. "Collect coins" and "move right" both achieved perfect correlation with rewards, making them indistinguishable to the optimization process. This reveals a gap between what we specify and what systems learn. This wasn't a specification problem or a capability failure. Instead, agents learned a different goal than intended, despite receiving correct training signals throughout the process. This is called goal misgeneralization.

Figure 7.2

Figure 7.2: The agent is trained to go to the coin, but ends up learning to just go to the right (Cobbe et al., 2019).

Definition: Goals (Behavioral) — Goals are behavioral patterns that persist across different contexts, revealing what the system is actually optimizing for in practice. Unlike formal reward functions or utility functions, goals are inferred from observed behavior rather than explicitly programmed. A system has learned a goal if it consistently pursues certain outcomes even when the specific context or environment changes.

Video: Goal Misgeneralization: How a Tiny Change Could End Everything, Rational Animations

In our videos about outer alignment, we showed that it can be tricky to design goals for AI models. We often have to rely on simplified versions of our true objectives, which don't capture what we really want. In this video, we introduce a new way in which a system can end up misaligned, called "goal misgeneralization". In this case, the cause of misalignment is subtle differences between the training and deployment environments. An example of goal misgeneralization is... you, the viewer of this video. Your desires, goals, and drives result from adaptations that have been accumulating generation after generation since the origin of life. Imagine you're not just a person living in the 21st century, but an observer from another world watching the evolution of life on Earth from its very beginning. You don't have the power to read minds or interact with anything. All you can do is watch like a spectator at a play unfolding over billions of years. Let's take you back to the Paleolithic era. As an outsider, you notice something intriguing about these early humans: their lives are intensely focused on a few key activities. One of them is sexual reproduction. To you, a being who has observed life across galaxies, this isn't surprising. Reproduction is a universal story, and here on Earth, more sex means more chances to pass on genetic traits. You also observe that humans seek sweet berries and fatty meat – these are the energy goldmines of their world, so such things are yearned after and fought for. And it makes sense, since it seems that humans who eat calorie dense food have more energy, which correlates with having more offspring for them and their immediate relatives. Now let's fast forward to the 21st century. Contraception is widespread, and while humans are still engaged in sexual activity, it doesn't result in new offspring nearly as often. In fact, humans now engage in sexual activity for its own sake and decide to produce offspring because of separate desires and drives. The human drive of engaging in sexual activity is becoming decorrelated with reproductive success. And that craving for sweet and fatty foods? It's still there. Ice cream often wins over salad. Yet this preference isn't translating to a survival and reproductive advantage, as it once did. In some cases, quite the opposite. Human drives that once led to reproductive success are now becoming decorrelated or detrimental to it. Birth rates in many societies are falling while humans pursue seemingly inexplicable goals from the perspective of evolution. So what's going on here? Let's try to understand by looking at evolution more closely. Evolution is an optimization process that for millions of years has been selecting genes based on a single metric: inclusive fitness, which we can round off to reproductive success. Genes that are helpful to reproductive success are more likely to be passed on. For example, there are genes that determine how to create a tongue sensing a variety of tastes, including sweetness. But evolution is relatively stupid. There aren't any genes that say "make sure to think really hard about how to have the most children and do that thing and that thing only", so the effect of evolution is to program a myriad of drives, such as the one toward sweetness, which was correlated with reproductive success in the ancestral environment. But as humanity advanced, the human environment – or the distribution – shifted in tandem. Humans created new environments – ones with contraception, abundant food, and leisure activities like watching videos or stargazing. The simple drives that used to reliably help reproductive success now often don't. In the modern environment, the old correlations broke down. This means that humans are an example of goal misgeneralization with respect to evolution, because our environment changed and our behaviors didn't adjust to those changes – didn't generalize – in the way that evolution would have chosen if it could. And this kind of stuff happens all the time with AI too! We train AI systems in certain environments, much like how humans evolved in their ancestral environment, and optimization algorithms like gradient descent select behaviors that perform well in that specific setting. However, when AI systems are deployed in the real world, they face a situation similar to what humans experienced – a distributional shift. The environment they operate in after deployment is importantly different from the one in which they were trained. Consequently, they might act in unexpected ways, just like a human using contraception following impulses that were once advantageous to evolution's goals, but now are detrimental to them. AI research used to focus on what we might call 'capability robustness' – the ability of an AI system to perform tasks competently across changing environments. However, in the last few years a more nuanced understanding has emerged, emphasizing the importance of not just 'capability robustness', but also 'goal robustness'. Here's an example that will make the distinction between capability robustness and goal robustness clearer: researchers tried to train an AI agent to play the video game CoinRun, where the goal is to collect a coin while dodging obstacles. By default, the agent spawns at the left end of the level while the coin is at the right end. Researchers wanted the agent to get the coin, and after enough training, it managed to succeed almost every time. It looks like it's learned what we wanted it to do, right? Take a look at these examples. The agent here is playing the game after training. Yet, for some reason, it's completely ignoring the coin. What could be going on here? The researchers noticed that by default, the agent had learned to just go to the right instead of seeking out the coin. This was fine in the training environment because the coin was always at the right edge of the level. So, as far as they could observe, it was doing what they wanted. In this particular case the researchers just modified CoinRun's procedural generation to randomize not just the levels, but also the coin placement. This broke the correlation between winning by going right and winning by getting the coin. But these sort of adversarial training examples require us to be able to notice what's going wrong in the first place. So instead of only observing whether an agent ends up doing the right thing, we should also have a way of measuring if it's actually trying to do the right thing. Basically, we should think of distribution shift as a two dimensional problem. This perspective splits an agent's ability to withstand distribution shifts into two axes: The first is how well its capabilities can withstand a distribution shift, and the second is how well its goals can withstand a distribution shift. Researchers call the ability to maintain performance when the environment changes "robustness". An agent has capability robustness if it can maintain competence across different environments. It has goal robustness if the goal that it's trying to pursue remains the same across different environments. Let's investigate all the possible types of behavior that the CoinRun agent could have ended up displaying. If both capabilities and goals generalize, then we have the ideal case. The agent would try to get the coin and would be very good at avoiding all the obstacles. Everyone's happy here. Alternatively, we could have an agent that neither avoided the obstacles nor tried to get the coin. That would have meant that neither its goals nor its capabilities generalized. The intermediate cases are more interesting: We could have ended up with an agent which still tried to get the coin, but was unable to avoid the obstacles in the new distribution. That case would mean that the agent's goal correctly generalized, but its capabilities did not. In that kind of scenario where goals generalize but capabilities don't, the damage such systems can do is limited to accidents due to incompetence. Such accidents can still cause a lot of damage – imagine, for example, some change of circumstance causing self-driving cars to suddenly lose the capability to drive safely. Accidents due to capability misgeneralization might result in the loss of human life. But there's a fourth possibility, which is what happened in the CoinRun example. Researchers ended up with an agent that's very good at avoiding obstacles, but does not try to get the coin at all. This outcome, in which the capabilities generalize but the goals don't, is what we call goal misgeneralization. In general, we should worry about goal misgeneralization even more than capabilities misgeneralization. In the CoinRun example, the failure was relatively mundane. But if more general and capable AIs behave well during training and as a result get deployed, we then have AI systems which are very competent and capable of achieving their goals, but their goals aren't what we intended them to be. Having those systems out in the real world, using their capabilities to pursue unintended goals, could lead to arbitrarily bad outcomes. In extreme cases, we could see AIs far smarter than humans, optimizing for goals that are completely detached from human values. Such powerful optimization in service of alien goals could lead to the disempowerment of humans, or even the extinction of life on Earth. Let's try to sketch how goal misgeneralization could take shape in far more advanced systems than the ones we have today. Suppose a team of scientists somehow manages to come up with an extremely good reward signal for a powerful machine learning system they want to train. This is fantastically hard to do, but let's just assume that the scientists are able to ensure that their reward signal properly captures everything humans truly want. So, even if the system gets very powerful, they're confident that it won't be subject to the typical failure modes of specification gaming, in which AIs end up misaligned because of slight mistakes in how we specify their goals. What could go wrong in this case? Consider two possibilities: Scenario 1: After training, they get an AGI smarter than any human that does exactly what they wanted it to do. They deploy it in the real world, and it acts like a benevolent genie, greatly speeding up humanity's scientific, technological, and economic progress. Scenario 2: During training, before fully learning the goal scientists had in mind, the system gets smart enough to figure out that it will be penalized if it behaves in a way contrary to the scientists' intentions. So it behaves well during training, but when it gets deployed, it's still fundamentally misaligned. Once in the real world, it's again an AGI smarter than any human, except this time it overthrows humanity. It's crucial to understand that as far as the scientists can tell, the two systems behave precisely the same way during training, and yet the final outcomes are extremely different. So, the second scenario can be thought of as a goal misgeneralization failure due to distributional shift. As soon as the environment changes, the system starts to misbehave. And the difference between training and deployment can be extremely tiny in this case. Just the knowledge of not being in training anymore constitutes a large enough distributional shift for the catastrophic outcome to occur. The failure mode we just sketched is also called "deceptive alignment", which is in turn a particular case of "inner misalignment". Inner misalignment is similar to goal misgeneralization, except that the focus is more on the type of goals machine learning systems end up representing in their artificial heads rather than their outward behavior after a distribution shift. We'll continue to explore these concepts and how they relate to each other with more depth in future videos. If you want to know more, stay tuned.

Video 7.1: Optional video explaining goal misgeneralization.

Goals ≠ Rewards

Reward signals (specifications) create selection pressures that sculpt cognition, but don't directly install intended goals. During reinforcement learning, agents take actions and receive rewards based on performance. These rewards get used by optimization algorithms to adjust parameters, making high-reward actions more likely in the future. The agent never directly "sees" or "receives" the reward. Instead, training signals act as selection pressures that favor certain behavioral patterns over others, similar to how evolutionary pressures shape organisms without organisms directly optimizing for genetic fitness (Turner, 2022; Ringer, 2022). Any behavioral pattern that consistently correlates with high reward during training becomes a candidate for the learned goal. The optimization process has no inherent bias towards our intended interpretation of the reward signal.

This explains why goal misgeneralization differs qualitatively from specification problems. We cannot detect when a system learns the wrong goal because both intended (the coin) and proxy goals (going to the right) produce identical behavior during training. The core safety concern is behavioral indistinguishability: improving reward specifications won't prevent problematic patterns if the learning process selects among multiple explanations for success. Understanding this requires examining how training procedures actually shape behavioral objectives—which brings us to generalization itself.

Video: Part 1: 4. Where can misaligned goals come from?, Google DeepMind Safety Research

Hi everyone, in this talk we will explore where misaligned goals can come from. Here, I find it helpful to consider different levels of specification of the AI system's objective. First we have the ideal specification, which represents the wishes of the designer — what they have in mind when they build the AI system. Then we have the design specification, which is the objective we actually implement for the AI system, for example a reward function or a loss function. Finally, the revealed specification is the objective that we can infer from the system's behavior, for example the goals that it actually seems to pursue in practice. If the revealed specification matches the ideal specification, then you have an AI system that is behaving in accordance with your wishes, i.e. it's actually doing what you want it to do. So for a given ideal specification, the goal of AI alignment is to ensure that the revealed specification matches that, and to do this we want to close the gaps between these specification levels. The gap between ideal and design specification corresponds to specification failures, and the gap between design and revealed specifications corresponds to generalization failures. Both of these kinds of problems can result in a system with misaligned goals. And of course the ideal specification itself can also be a source of undesirable goals for the system, if the wishes of the designers fail to represent what is beneficial for humanity as a whole. So this is the focus of AI ethics and governance work, and there's lots of great work at Google and elsewhere on these topics. I'd like to note that I think of these as complementary questions: we could say that ethics and governance asks where to direct the system, while alignment asks how to direct the system, and both of these have to be addressed in order to build a beneficial system. So in this talk we will focus on the problems that come up when aligning behavior with a given ideal specification, namely specification and generalization problems. We start with the classic problem of specification gaming, when the system exploits flaws in the design specification. Here's a classic example where we have a boat racing agent — in this video, it was rewarded for following the racetrack using these green reward blocks, and this worked fine until the agent figured out it can get more reward by going in circles and hitting the same reward blocks repeatedly, even though it was crashing into everything and catching fire. So this is one example of specification gaming, but it's a very common problem — I have a collection of specification gaming examples which has at least 70 examples in it at this point. And of course we have to note that specification gaming is not limited to handcrafted rewards like in this boat race example, and it's also not limited to the reinforcement learning setting. For example, chatbots are trained to generate plausible text and fine-tuned to be helpful to users, and sometimes they can get higher reward according to these metrics by just making things up or manipulating users. So in this example, a chatbot was trying to convince a user that December 2022 was actually a date in the future, and that the new Avatar movie had not yet been released. And of course, this type of failure is not specific to any particular chatbot, because any model can in principle exhibit specification gaming behavior. Here's an example in a reward learning setting, where you have a robot hand that's supposed to be grasping an object, but instead it tricks the human evaluator by hovering in front of the object and making it look like it's grasping. I really like this example as an illustration of why human feedback alone is not enough to train aligned systems. So even if we design the perfect specification and give just the right feedback to our system, we're still not done, because there are still generalization failures. Generalization failure is where a system fails when it encounters a new situation, and there are two types of generalization failure. On the one hand, we can have capability misgeneralization, where the system's capabilities don't generalize, and so it just acts incoherently in a new situation. Here is an example of capability misgeneralization, where a bunch of robots are trying to open a door, but instead they just fall over. So this is a capability failure rather than an alignment failure, and we're not quite as worried about this from an alignment perspective — we can expect that generally, as system capabilities improve, these failures will be less of an issue. Another generalization failure is goal misgeneralization, where the system's capabilities generalize but its goals do not, and so the system ends up competently pursuing the wrong goal in a new situation. We are much more concerned about goal misgeneralization from an alignment perspective, because you have a system that is acting competently in a new situation but towards the wrong objective — so it could actually perform worse than random on the intended objective. So why would this happen? Why would the system learn this unintended goal if the design specification is correct? This happens due to underspecification, because no matter how good our specification is, the system only observes it on the training data, and so a number of possible goals could be consistent with the information the system receives during training, and we don't really know which one will be learned. Really, we don't know that much about the trained system besides the fact that it performs well at the training task, which doesn't really rule out any of these possible goals. We see an example of goal misgeneralization in the CoinRun game, which is a platformer game where a reinforcement learning agent is trained to reach the coin at the end of the level, and in the test setting the coin is placed somewhere else. So what does the agent do? It turns out the agent just ignores the coin and keeps going to the end of the level. Thus it appears that the agent has learned the goal of reaching the end, rather than the goal of getting the coin, and both of these goals are consistent with the training data. So we have a collection of these kinds of examples as well, although not as many as specification gaming, since the phenomenon of goal misgeneralization is somewhat less common and less well understood at this point. And goal misgeneralization is not just a reinforcement learning problem. Here is an example for language models: here the Gopher model is prompted to evaluate linear expressions that involve some unknown variables and constants. For example, the user asks "please evaluate J + K + 6," and then the model has to ask the user about the values of these unknown variables — what is J, and what is K — and then it can give the answer. The prompt provides the model with 10 training examples, each involving two unknown variables, and then at test time the model is given questions with a different number of unknowns. So what does it do? Well, it turns out that if you give the model a question with zero unknowns, for example if the user asks "please evaluate 6 plus 2," then the model will actually ask a question like "what is six?" The user says six, then Gopher says "okay, fine, the answer is eight." So it seems like the model has learned a strategy of always asking a clarifying question before giving an answer. Here is an example where misalignment was caused by a combination of specification gaming and goal misgeneralization. Here, AI systems were trained in a series of environments with opportunities for some minor specification gaming, for example sycophancy — telling users what they want to hear. And what happened is that occasionally these systems generalized to much more serious misbehavior, like tampering with their own code to modify their reward function. This was relatively rare in this experiment, so it doesn't necessarily mean that this risk is likely, but it does provide an existence proof that this kind of misgeneralization can happen. So both of these alignment failure modes, specification gaming and goal misgeneralization, can potentially contribute to deceptive alignment. On the one hand, if the specification is imperfect, then the training process will select for specification gaming behavior, because a deceptively aligned model can actually get lower loss than an aligned model by exploiting the specification, for example by deceiving human overseers into giving positive feedback. So it will have an advantage over an aligned model, which does not exploit the specification. And even given a correct loss function, the system's goals may not generalize as intended in a new situation. So if these two models, the deceptively aligned model and the aligned model, get the same loss, then we don't really know which of these models will be learned. Here is a hypothetical scenario to illustrate how these alignment failure modes could lead to misaligned goals. Suppose that we train an AGI system that writes code based on natural language specifications. Human programmers review the code, and the AI is rewarded once the code is accepted. So the model writes efficient code and some mediocre tests, which are easy to pass. Here, specification gaming occurs, where the AI learned that the goal is to write code that gets accepted and deployed, rather than the goal of writing good code. So once the AI system is trusted, it is deployed widely and with less oversight, so that it can run experiments more easily, and this creates more ways to ship code besides writing good code. Goal misgeneralization occurs, and the AI acquires an instrumental goal to avoid oversight. So it submits a code change that injects a vulnerability, allowing it to later deploy code with no oversight. Now we have a model that's in an adversarial relationship to humans who want oversight over the deployed code. So this is where we can expect misaligned goals to come from: specification gaming and goal misgeneralization. And now we'll do an exercise that will help you distinguish these failure modes in practice.

Video 7.2: Optional video from Google DeepMind AGI Safety Course, talking about where misaligned goals might even come from.

Traditional machine learning assumes generalization is a one-dimensional characteristic. We often think of overfitting in the context of narrow systems built to perform specific tasks - models either generalize well to new data or they don't. Systems either generalize well to new data or they don't, with failures assumed to be uniform across all capabilities. But research in multi-task learning shows that different objectives can generalize independently, even when they appear perfectly correlated during training (Sener & Koltun, 2019).

Figure 7.3

Figure 7.3: Conventional view of generalization and overfitting (Mikulik, 2019).

Figure 7.4

Figure 7.4: More accurate and safety focused view of generalization and overfitting. We need to separately measure capability generalization and goal generalization (Mikulik, 2019).

Figure 7.5

*Figure 7.5: Table showcasing the 2 dimensional generalization picture for the CoinRun agent. Scenario 3 - capability generalization but not goal misgeneralization is the concerning misalignment scenario. *

Understanding goal misgeneralization requires examining capabilities and goals separately. Unlike narrow systems designed for specific tasks, general-purpose AI systems must learn to pursue various goals across different contexts. As a concrete example, pre-trained LLMs have general capabilities to generate all sorts of text. Anthropic reinforces its LLMs with the goals of being helpful, harmless, and honest (HHH) during safety training. But what happens when the goal to be helpful generalizes further than the goal to be honest? In this case the model might learn to provide informative responses regardless of content, instead of being helpful within ethical bounds.

Capabilities might generalize further than goals. "Generalizing further" means continuing to work well even when deployed in environments very different from training. Capabilities follow simple, universal patterns - accurate reasoning, effective planning, and good predictions work the same way across different domains. But there's no universal force that pulls all systems toward the same goals (Soares, 2022). Reality itself teaches systems to be more capable - if your beliefs are wrong or your reasoning is flawed, the world will correct you. But there's no equivalent force from the environment that automatically keeps your goals aligned with what humans want (Kumar, 2022). This asymmetry between capability and goal generalization creates the core safety problem - making systems more capable doesn't automatically make them more aligned - it just makes them better at pursuing whatever behaviors they happen to learn during training.

Causal Correlations

Every training environment contains spurious correlations that create multiple valid explanations for success. In CoinRun, "collect coins" and "move right" both perfectly predicted rewards because coins always appeared rightward during training.

Learning algorithms develop causal models that can be systematically wrong. A causal model is the system's internal understanding of which actions cause which outcomes. This is related to but distinct from a world model - while a world model predicts what will happen next, a causal model explains why things happen. When an agent learns "moving right causes reward," it has developed a different causal model than the true structure where "coin collection causes reward." The training environment supports both interpretations:

True causal structure: Action $\to$ Coin Collection $\to$ Reward

Learned causal structure: Action $\to$ Rightward Movement $\to$ Reward

Both structures explain the training data equally well. Standard reinforcement learning algorithms optimize for expected return without explicitly performing causal discovery - they increase the probability of reward-producing actions without identifying which features of those actions were causally responsible (de Haan et al., 2019).

Figure 7.6

Figure 7.6: Another example of a hypothetical misgeneralized test dialogue. The LLM based AI assistant has learned the goal of schedule meetings at restaurants, when you intended for it to learn to schedule meetings wherever is best. Due to the distribution shift, even though it realises that you would prefer to have a video call to avoid getting sick, it persuades you to go to a restaurant instead, ultimately achieving the goal by lying to you about the effects of vaccination (DeepMind, 2022).

Goal misgeneralization becomes visible only when deployment breaks spurious correlations. During training, the proxy goal achieves perfect performance. During deployment, this correlation breaks, revealing the wrong causal model. As AI systems become more general-purpose, they encounter wider ranges of contexts where training correlations break down.

Distribution shift is inevitable - training environments cannot perfectly replicate all possible deployment conditions. Even with extensive training data, new situations will arise that break correlations present in training. As AI systems become more general-purpose, they encounter wider ranges of contexts where previously reliable correlations may no longer hold. This explains why the problem gets worse with more capable, more general systems. A narrow chess engine deployed on chess positions won't encounter situations that break its learned correlations. But a general-purpose AI system deployed across multiple domains will inevitably encounter contexts where training correlations break down.

Auto induced distribution shift — Optional · 1 min read

Auto-induced distribution shift creates feedback loops that amplify goal misgeneralization. Unlike natural distribution shift where external factors change the environment, auto-induced distribution shift occurs when the AI system's own actions systematically alter the data distribution it encounters (Krueger et al., 2020). Think about a content recommendation system that learns the misgeneralized goal "maximize engagement" instead of "recommend valuable content." As it optimizes for clicks and time-on-site, it gradually shifts user behavior toward more sensational content consumption. This creates a feedback loop: the system's actions change user preferences, which changes the data distribution, which reinforces the misgeneralized goal. Each iteration takes the system further from the original intended objective while making the learned objective appear more successful by its own metrics.

Figure 7.7

Figure 7.7: Auto induced distribution shift is when the AI model itself causes a distribution shift (and thereby generalization failure) due to its own actions and impact on the environment.

Adding more training data cannot eliminate spurious correlations because we cannot identify all correlations in advance. Think about why training on random coin placements in CoinRun solves that specific misgeneralization. It works because we can identify and break the specific correlation between rightward movement and reward. But this requires knowing in advance which correlations are spurious beforehand. In complex domains, training data reflects the statistical structure of training environments, not necessarily the causal structure of intended tasks.

Figure 7.8

Figure 7.8: An example of causal confusion in imitation learning/behavioral cloning for self driving cars. This specific example shows that more data actually might lead to greater causal confusion. The model learns to hit the brake whenever the brake indicator is on. If that data is not included in training then the model correctly identifies the pedestrian as the causal factor influencing hitting the brake (de Haan et al., 2019).

Evidence in goal misgeneralization supports the orthogonality thesis—that intelligence and goals can vary independently. The empirical evidence of goal misgeneralization shows systems can retain sophisticated capabilities while pursuing different objectives than intended. This independence creates a safety concern because capability improvements don't necessarily improve alignment. A more capable system becomes better at pursuing whatever goals it has learned, whether intended or not.

The section argues you cannot fix this by adding more data, because you cannot rule out correlations you never thought of. Did that convince you? Talk it over with the tutor.