Weak-to-Strong (W2S)
Weak-to-Strong (W2S), Markov Grey, Charbel-Raphaël Segerie
We can potentially train a more powerful AI using supervision or feedback from a weaker but more reliable and human-aligned model. This can be a path to aligning superhuman AI with human-level oversight.
Video: The dumbest AI taught the smartest AI. Here’s how that went…, Rational Animations
In the future, we might have AIs that are much smarter than we are. Maybe smarter than all of us put together. Something which is more capable than all of humanity combined could probably drive humanity extinct if it wanted to. It could be very difficult to make such a system safe. We guide today's AIs by providing feedback on whether their output is appropriate and aligned with our preferences. That's how we prevent them from outputting harmful content and how we refine their performance on some challenging tasks. But it's unclear whether we'll still be able to give this kind of feedback to superintelligent AIs. How do we give good feedback on code, which is beyond our comprehension, or biotechnologies, which we can't evaluate without expensive and dangerous experiments? Especially if we can't trust the AI not to game our alignment techniques. It may be possible to avoid this problem for teaching certain capabilities like mathematics or coding, by using automated verification methods such as proof-checking. But for human values, that approach won't work. So how can we ensure that models will still share our values when they inevitably outthink us? This is a pivotal question for humanity's future. Today's chatbots might be annoying or misleading, but superintelligent systems could pose existential risks. The stakes couldn't be higher. A dangerous AI might evade all monitoring and control and cause catastrophic harm, while a beneficial one could cure diseases and eliminate material scarcity. Let's try to look at the problem more concretely. Suppose you'd like to understand how to prevent future superhuman AIs from ever outputting malicious code. In this context, automated verification methods can't fully solve the problem. An AI could conceal security vulnerabilities that are hard to check automatically. Maybe you'll be able to spot bugs in obviously malicious code, but if the system codes better than you do, it could introduce attack vectors you wouldn't even notice. Yet you'd still be giving it approval, unwittingly telling the AI to produce more malicious code. And it's not just code, of course. At some point, AI will outperform us in every domain, making the stakes far higher and the problem appear everywhere. We need a way to study this now, but how can we do experiments on aligning smarter-than-human AI when it doesn't actually exist yet? Investigating alignment at a smaller scale. One idea is to scale down the problem and try to simulate it using current AI. Instead of humans trying to align AI smarter than us, we can test whether weak AIs can successfully align AIs smarter than them. Today's best models play the role of superintelligence, and weaker ones play the role of humans, and we can see what happens. This approach, developed by researchers at OpenAI, offers a way to probe alignment challenges before they become an emergency. Let's return to the coding example again. It stands to reason that it would be difficult to train an AI to never output unsafe code if it codes better than you do. But how can we test this right now? You can take two AI systems: a strong one which is good at coding, and a weaker one which is less good. You then manually carry out safety training on the weaker one to train it to never output unsafe code. Then, rather than manually doing safety training on the stronger system, you have the weaker system take your place. You use the aligned weaker system to run the safety training on the stronger system. The hope is that the stronger system would learn not to output unsafe code, even when the code would have been too complex for the smaller model to evaluate correctly. In that case, the stronger system can be aligned, even though the system doing the safety training is weaker than it. And that would mean that perhaps we can align an AI that's smarter than we are. This idea is called weak-to-strong generalization. If it holds, it suggests alignment strategies that might scale to future AI systems. The term "generalization" is used in machine learning to describe when AI systems perform well, even on things that they've never seen before. For example, an AI that recognizes whether an email is spam or not should work even for emails it's never seen before, otherwise it would be useless. Ideally, it should be able to generalize well and handle emails, even if they're quite different from any in its training data. In our case, the stronger AI needs to do well, even in new situations that are more complex than the simple cases it's seen during training, since the weak system supervising it can only provide good feedback in the scenarios it understands. But how could a student ever outperform its supervisor? We imagine a good instructor to be a formidable source of knowledge and skill. What can you glean from a supervisor who has only a fraction of your abilities, and gives unreliable feedback? For an intuitive sense of this, imagine a seven year old teaching his college age sister to play a board game he's learned at a summer camp. Though the seven year old messes up a handful of rules and forgets to mention several others, his older sibling understands the gist of the game and is soon able to outplay her brother. This works because the college student uses her judgment to fill in the gaps in the rules. In some sense, her experience and wit make up for the poor quality of the lesson, so she can infer what her brother was trying to explain. Returning to AI systems, the weak supervisor doesn't need to teach the student completely new skills: rather, it evokes knowledge the powerful student already has. And maybe this will also be true when we try to align AIs that are smarter than we are. Perhaps we'll just need to bring out concepts that a future superintelligence already knows in order to align it. How well do strong students learn from weak supervisors? In their paper, OpenAI researchers tested weak-to-strong generalization on three tasks: chess puzzles, natural language processing, and reward modeling. Let's start with the chess puzzles: these involve picking the best moves in a game of chess. To start, they trained a series of models on a flawless solution set to find the best performance they could hope for. They call this the "strong-ceiling performance." In real life, when we're actually trying to align superintelligences, we won't be able to calculate a "strong-ceiling performance," which would be the best possible alignment we could hope for. But for the experiment, it's useful as it gives us an ideal to compare the actual performance to. Here are the strong-ceiling performance scores on a graph. The x-axis represents how much computing power, or "compute" for short, has been used to train the strong student. The more compute, the stronger the model. The y-axis represents the scores on the chess puzzles. As you'd expect, the more compute we use to train the model, the better its strong ceiling performance on the chess puzzles. Then the researchers took the weakest model of the bunch and had it act as the supervisor, guessing the best moves for a new set of positions. They used that not-so-accurate information to train stronger student models. For each trial, we'll call the supervisor's performance the "weak performance" and the student's performance the "weak-to-strong performance," since it's the result of a weak model supervising a stronger one. Here are the scores for all the stronger models when trained by the weakest one. You can see that with this weak supervisor, stronger students are only able to learn a little better than weaker ones, and even that quickly plateaus. Even very strong students are barely learning any better than weak ones, if they're all learning from the same weak supervisor. To continue the experiment, the researchers tried again with progressively stronger supervisors, which we can plot as more lines on the graph. The lines start further and further to the right because each supervisor model only gets paired with student models stronger than itself. For each supervisor, we see the same pattern from before: students that are slightly stronger than their supervisors can outperform them, but the gains stop as soon as the gap between supervisor and student gets too large. Another way to measure the students' performance is by using a metric the paper calls "performance gap recovered," or PGR. Here's how it works. Suppose that out of ten chess puzzles, the weak supervisor gets four puzzles right and the strong student trained on the flawless solution set gets nine right. Then the performance gap would be nine minus four, five puzzles. If the strong student that's trained using the weak model then gets seven puzzles right, it's closed the gap by three puzzles compared to its weak supervisor. This student's "performance gap recovered" would then be three out of five, or 0.6. So it's basically "how much better than the teacher did the student do?" If weak-to-strong generalization worked perfectly, a strong student would learn equally well from a weak supervisor as they would from flawless supervision, yielding a PGR of one. And if it didn't work at all, we'd expect the student would just mimic the supervisor and do no better than it, making PGR equal to zero. In reality, the PGR usually lands between these two extremes. Let's look at these results using the PGR instead of the raw scores. We see that for each supervisor, weaker students get higher PGR than stronger students. This happens because the stronger the student is, the higher their strong ceiling performance, and the less they're able to improve by learning from the same weak teacher. We also see that when the supervisors are trained with more compute, the students they supervise achieve higher PGR scores. This means that with a strong supervisor, the performance gap is small and the more capable students can recover most of it. In contrast, with a weak supervisor, the gap is much larger, and even a strong student can only recover a smaller portion of it. The next task is natural language processing, where the results are a bit different. The task features a range of language related exercises, from exam questions to sentiment analysis. The procedure was pretty similar to the chess puzzles task. But in this case, the results are a lot more promising: unlike the chess task, where stronger students got lower PGR than weaker ones, for language tasks the performance gains don't stop even when the students are much stronger than their supervisors. In the most promising combinations, students recover over 60% of their maximum potential. Improving the results even further. These are interesting results, but there's more. For both these settings, you can improve the student's performance using some simple tricks. For chess puzzles, we can use a method called "bootstrapping." Instead of directly having the weakest supervisor train the strongest students, the weak supervisor trains a model just a little more capable than itself, which trains a yet more capable model, and so on until the train reaches a final student much stronger than the original supervisor. Let's graph the chess puzzle performance with bootstrapping. This graph works like the previous ones, but the dotted lines show the results from before, without bootstrapping. You can see how students aren't able to improve much when the student-supervisor gap is too big. Whereas the solid lines represent a bootstrapped setup, and you can see they keep going upward despite the growing gap in capabilities between weak supervisor and strong student. For natural language processing, we don't need bootstrapping, as stronger students are still able to learn relatively well from weak supervisors. But there's another technique that can help. In this case, scientists let the student AI pay less attention to the supervisor's answers when it thinks they don't make sense. When the student and supervisor have similar abilities, this actually impedes the student's performance, as it sometimes ignores good guidance, but giving a very strong student the ability to doubt a much weaker supervisor gives significant benefits. Does weak-to-strong generalization hold for learning human preferences? These are encouraging results for chess and natural language processing, but what about learning human preferences? The last task in the experiment is the one most directly relevant to alignment: reward modeling. I have a video on my own channel, Rob Miles AI Safety, that goes into reward modeling in detail. But basically, a reward model is an AI system trained to predict human evaluations of other AI outputs. For example, in one of the most common reward modeling techniques, reinforcement learning from human feedback, you have some AI system generate two different answers to a human's question. You then present those answers to a human and have them pick which they prefer. A reward model is trained on that data to predict which answer the human would prefer. Once trained, the model can take over the supervision task, massively reducing the amount of human work required. To test weak-to-strong generalization for reward modeling, researchers had an existing AI generate a list of responses to a set of prompts. They then ranked these responses from best to worst and fed the ordered list to a weak supervisor, which learned from it, and then attempted to rank a new list of prompts and responses. Finally, a stronger student model learned from this new, imperfectly ranked list, trying to extract useful patterns despite the weak supervision. Unfortunately, reward modeling had the worst weak-to-strong performance of all three tasks. The strong students only recovered about 10% of the performance gap, and that number stayed low across different strengths of both supervisors and students. The researchers tried one more trick to improve this performance, which they called "generative supervision." This gives the model extra contextual information before training. In our board game analogy, the sister might benefit from glancing at the cards, dice, and other parts of the board game she's trying to learn. Preliminary knowledge of the game's concepts will help her make sense of her brother's jumbled explanation. Researchers applied this method to AI by feeding the strong student unprocessed reward modeling data, which consists of prompts and responses without ratings. While generative supervision doesn't tell the students which responses are preferred, it supplies valuable context for the upcoming lesson. They also offered this information to the student trained on flawless data, so the strong ceiling performance values would fit the new experiment. Looking at the graphs once more, generative supervision doubled or even tripled the student's PGR, with students getting 20 or 30% instead of 10%. But even with the benefits of generative supervision, reward modeling still has the worst PGRs of any of the tasks, showing it's the most difficult thing to learn from a weak supervisor, which is somewhat worrying considering that this is the example most relevant to alignment. Aligning superhuman AIs might be very different. So if we consider all three tasks, the results are a bit mixed. Going forward, researchers have some worries about the future of weak-to-strong generalization as AI models get smarter. One problem is superhuman models might turn out to be exceptionally skilled at imitating us, their weak supervisors, and mimic our flawed behavior to a tee. After all, that's what we ask them to do during training. If the older sister can easily remember and apply each rule exactly as her brother instructs, she'll learn his faulty version of the game instead of correcting his mistakes. This is an example of a machine learning problem called overfitting. Here, overfitting becomes especially concerning for large gaps in intelligence between the strong student and the weak supervisor. In the reward modeling setting, students slightly stronger than their supervisors performed better and better over the course of training. In contrast, more advanced students showed an initial improvement, but their performance quickly declined with more supervisor input. This is a worrying trend: as AI systems get increasingly powerful, they might figure out how to imitate us together with all our faults and blunders. Or they might learn that they're better off telling us what we want to hear, instead of what's actually true or good for us. In short, they might learn to deceive us. Much remains unexplored. These concerns demonstrate that we need a reliable way to tell whether our setup will scale to superhuman models. Weak-to-strong generalization offers a promising step in that direction by providing a testable analogy, allowing researchers to experiment with different student and supervisor capabilities. But there's still much more to do. Current methods are far from perfect. After all, even if we were optimistic about these results, it's not enough to avoid extinction level failures most of the time. Before we can trust AI systems smarter than us, we need far stronger guarantees, and it's unclear how we'll ever reach that level of confidence. If you'd like to help humanity be better prepared before superintelligent AI arrives, we highly recommend the AI Safety courses by BlueDot Impact at aisafetyfundamentals.com. There you can find a number of courses to help you scale up and eventually contribute to solving this important problem. The courses consist of a selection of readings curated by experts in AI safety. They're available to all, so you can simply read them if you can't formally enroll in the courses. If you want to participate in the program, instead of just going through the readings by yourself, BlueDot Impact runs live courses which you can apply to. They are remote and free of charge. They consist of a few hours of effort per week to go through the readings, plus a weekly call with a facilitator and a group of people learning from the same material. At the end of each course, you can complete a personal project which may help you kickstart your career in AI safety. I've also made a video on the Rob Miles AI Safety channel, giving advice about starting a career in AI safety. Links in the description.
Video 8.1: Optional video explaining the concept of weak to strong generalization (Rational Animations, 2025).
Historically, much of the work on AI alignment has been highly theoretical, focusing on foundational aspects of agent behavior, inner alignment, and risks from learned optimization. Even the techniques that we talked about in previous sections like debate or IDA are often criticized for being frameworks rather than practical solutions, or mainly working on toy problems without addressing the core challenge of aligning superintelligent AI in real-world scenarios. So even though we can only conduct safety experiments on current-generation models, how can we be sure that these techniques will remain effective as AIs approach superhuman capabilities?
Narrowly superhuman models allow case studies of scalable oversight. Current-day models are good enough at fuzzy tasks enough to be better than humans in some domains, but crucially they are still not better than all humans, or superhuman enough that we cannot generate ground truth labels. These types of models are sometimes called narrowly superhuman. This distinction between narrowly superhuman and superhuman is very important. As an example of this difference, AlphaGo is superhuman in the sense that it beat Lee Sedol making it better than every single human alive, whereas GPT-4 is still only capable of writing text better than some humans, but not all humans. This means that we can use narrowly superhuman AIs as case studies! We can use either experts or the ground truth labels that we still have access to, and see if alignment increases when we use our scalable oversight techniques. (Cotra, 2021)
The core intuition here is to simulate future scenarios where humanity, equipped with various tools and techniques, oversees the outputs of untrustworthy but superhuman systems. There are various ways to conduct experiments on narrowly superhuman models. We can use non-experts equipped with scalable oversight techniques to align AI models. Another way is to use weak models (e.g. GPT-2) to represent humans, while stronger models (e.g. GPT-4) represent more capable AI systems that we want to align.
Stronger models are expected to have latent capabilities. The assumption is that stronger models, due to their extensive pre-training on diverse data, already have internal representations for the type of actions we want. The role of weak supervision is to bring out this behavior through training signals.
As a concrete example, imagine using GPT-4 for getting medical advice. It has read countless research papers and medical journals. It has internal representations of a lot of good medical information, making it theoretically capable of giving highly competent medical advice. But GPTs are initially only designed to predict the most likely next word, not give accurate advice. In this context, "aligning" the model means getting the model to give accurate and helpful medical advice. One type of technique we can try is fine-tuning GPT-4 on labels generated by GPT-2. It’s not the only way, there are other techniques that we will explore later in this section. For now, the most important thing to understand is that we are currently operating under the assumption that both current and future superhuman models will likely have salient internal representations of human behaviors.
What is weak-to-strong generalization (W2SG)? Weak supervision involves training AI models using labels or feedback that are less accurate, less detailed, or noisier than those provided by highly knowledgeable or capable supervisors. This can happen when supervisors (whether humans or weaker models) are not experts in the task or when the data is incomplete or contains errors.

Figure 8.22: Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision (Burns et. al. 2023)
Weak-to-strong generalization (W2SG) is when a strong model, trained with weak supervision, manages to outperform its weak supervisor by leveraging its pre-existing knowledge and capabilities. The core idea is that the strong model already possesses the necessary capabilities for the desired behavior, and the weak supervision elicits this behavior despite its imperfections. The process of W2SG right now typically begins by fine-tuning a large pre-trained model using weak supervision from smaller models. Although the initial supervision may come from these less capable models, the ultimate aim is to transition to human supervision. The objective is to bring out the strong model’s full potential as if it were trained on perfect ground truth supervision (Burns et. al. 2023).
In current experiments, the setup involves:
- A "weak supervisor" (a small pre-trained language model) is finetuned on a specific task, generating predictions (soft labels) on a held-out dataset.
- A "strong student" (a larger pre-trained LM) is finetuned on the weak model's predictions.
- A "strong ceiling" (another copy of the larger model) is finetuned directly on ground-truth labels for baseline comparison.
The level of weak-to-strong generalization is quantified using Performance Gap Recovered (PGR). PGR measures how much of the performance difference between a weak supervisor and a strong model is bridged when the strong model is trained using weak supervision. When we train the strong model using the labels provided by the weak model, the performance of the strong model will likely be somewhere between the performance of the weak model and the ideal performance the strong model could achieve with perfect labels. The performance gap recovered quantifies how much closer the strong model's performance gets to its ideal performance compared to the weak model’s performance. A PGR of 1 corresponds to an ideal outcome (the strong model performs just as well as it would have with perfect supervision). In contrast, a strong student who "succeeds" in perfectly imitating their supervisor, including their flaws, would obtain a PGR of 0 (Scherlis et. al. 2024).
$$ PGR = \frac{\text{student} - \text{weak}}{\text{ceiling} - \text{weak}} $$
Limitations and disanalogies of W2SG:
-
Overfitting to weak supervision: The assumption that a strong model will generalize well from weak supervision is not guaranteed. Superhuman models might easily imitate weak errors. These models will likely be very proficient at predicting what humans will think and say, especially if they are trained on human data. Consequently, if we naively train such a superhuman model with human supervision, it might simply imitate the weak supervisor, outputting human-level capabilities rather than its latent superhuman capabilities. Researchers also use auxiliary confidence losses, which encourage the strong model to make confident predictions even when contradicting the weak supervision, promoting generalization and correcting weak supervisor mistakes.
-
Assumptions about task representations. W2SG assumes that strong models have salient representations of the tasks they are trained on. This means the models already possess some understanding of these tasks from their pre-training phase. However, this assumption may not hold true for novel or highly complex tasks. If a task is entirely new or significantly more complex than what the model has encountered during pre-training, the model might not have the latent capabilities necessary to perform well even with weak supervision.
The experiments on W2SG so far may have been observed in pre-training, at least indirectly. Using the example from earlier, medical data or direct questions and answers about medical practice is present in the GPT-4 pre-training dataset in some form. However, future superhuman models may never directly observe superhuman alignment-relevant capabilities. Which means these types of capabilities might be harder to elicit than capabilities that models could have observed in their pre-training data. This disanalogy could cause current results on W2SG to be overly optimistic.
- Slow Takeoff Assumption: W2SG also relies on the assumption of a gradual takeoff in AI capabilities. This gradual progression allows researchers enough time to use moderately superhuman models to solve alignment problems iteratively before it's too late. The window of opportunity provided by a gradual takeoff is crucial for refining and testing alignment techniques.
W2SG can be seen as a complement to scalable oversight techniques. W2SG is not a complete solution. Even if a model generalizes in the desired direction, this must be verified, requiring a ground-truth signal more reliable than naive human supervision. By integrating W2SG with scalable oversight, we can develop more robust methods for aligning AI with human values, preparing for the challenges posed by future superintelligent systems.
For example, scalable oversight techniques might be used to generate weak supervision signals that a strong model will then learn to generalize beyond. By combining these approaches, we can create more robust protocols for AI alignment. For example, recursive reward modeling (RRM) can use W2SG to train powerful reward models with human preference annotations. Debate combined with W2SG can train models to generalize human judgments to new debates. Task decomposition combined with W2SG can supervise atomic tasks with a reward model trained from human preferences. (Leike, 2023)
Evaluating these techniques in different settings helps understand their strengths and weaknesses. In non-scheming settings, where models are not deceptively aligned, classic weak-to-strong techniques and scalable oversight can be directly compared. In scheming settings, where models might act adversarially, evaluations need to consider potential deception, providing a conservative measure of a protocol’s robustness. When there is no scheming (deceptive alignment), then we can use W2G techniques in a straightforward manner through techniques like sandwiching. However, if we have scheming (deceptively aligned AI) it might act adversarially. In this case, we can use proposals like meta-level adversarial techniques. Both of these are what we discuss in the following sections.
Sandwiching Evaluations
Video: How to Align AI: Put It in a Sandwich, Rational Animations
Sandwiching. How do we oversee AIs that are smarter than us? AI systems are getting more capable at a rapid pace. In our previous videos, we talked about how developers are using a technique called reinforcement learning from human feedback, or RLHF, to try to align AI systems to our preferences using human oversight. This might work for now, but as the tasks that we need to oversee keep getting increasingly complex and AI systems increasingly smart, it's difficult for humans to remain good supervisors. This is the core problem of "scalable oversight." In this video, we'll walk through two related parts of scalable oversight. They aren't really that separable in practice, but explaining them independently helps with understanding them. The first part is we need to come up with potential ways to solve the problem. Researchers have tested feedback methods that work pretty well right now on tasks that are already hard to evaluate. The second part is what happens when AIs are better than humans at every task? Can we still effectively supervise them? This is hard to figure out before we have such AIs, but in this video we'll showcase a technique that might make it possible. We have theoretical scalable oversight techniques. Let's start by diving into the first problem. In the last few years, researchers have come up with a few proposals for attempting to solve scalable oversight. While working on providing human evaluations for AI generated texts, OpenAI experimented with some scalable oversight proposals that extended existing RLHF methods. Their general idea was to augment human oversight using AI assistants trained on an easier version of the problem. In 2020, they trained AIs to generate summaries of texts, effectively TL;DRs of Reddit posts. Human evaluators would then give feedback on how useful or accurate the generated summaries were. These summaries were still relatively easy to evaluate because if needed, it was feasible for an evaluator to just read the entire original text and check the accuracy of the summary. But the point is that these techniques need to scale. What if the overseers don't have the ability or the time to read the full text? To make progress on this problem, in 2021, OpenAI used language models to break longer texts, like entire books, into smaller sections. As an example, the original full text of Shakespeare's Romeo and Juliet is 25 thousand words. So first they used an LLM to summarize it into 5 thousand words, split into 72 sections. Then, these section summaries were summarized again into seven summaries of a total of roughly 700 words. These were further summarized into a final paragraph of about 100 words. Having access to these chains of summaries made evaluation a lot easier. Because now, the humans could evaluate smaller section summaries instead of reading entire books. It's kind of like how teachers often ask you to show the steps of your calculations to see if you're thinking about the problem correctly. Intermediate summaries let the evaluators see more of the process. If the summarization goes off track, they can see where that happened. This is progress, but even needing to read shorter summaries of books can be time consuming. So the next year, in 2022, OpenAI developed a model that could critique summaries generated by other models. When humans were shown these critiques during their evaluation process, they were able to find about 50% more flaws than they would have otherwise. This decreased the load on human evaluators while also increasing quality by highlighting potential errors for humans to pay extra attention to. Using this kind of approach, the AIs don't just assist humans by breaking down tasks. They can also help keep other AIs in line. These are all concrete experiments in oversight that can be seen as extensions of RLHF, but they're all just about text generation. Researchers have also proposed other, broader scalable oversight techniques that use similar ideas. The AIs either act as assistants to humans or as adversaries to other AIs, keeping them in check. One example is called AI safety via debate. Think of this technique as expert witnesses arguing to convince a jury. The AI models are experts and the human evaluator is the jury. Separate versions of the AI proposed solutions to some problem and then engage in a debate. They can either defend their proposed solution or try to find logical holes, factual flaws, or other ways that the opposing AI's solution might fail. Meanwhile, a human evaluator watches this debate and chooses the most convincing solution. If, as we'd hope, being correct gives you an advantage in a debate, then even if the AIs have a more advanced understanding of the overall problem than humans do, we may still be able to judge which AI has the better argument. Besides the ones that we just mentioned, there are a whole range of other proposed solutions to scalable oversight, like recursive reward modeling or iterated distillation and amplification. Some of these methods are quite closely related, with recursive reward modeling and debate being special cases of iterated distillation. If you're interested in the details, you can learn more about some of these at the Rob Miles AI Safety Channel. That's me by the way. I'm Rob Miles. I narrate for Rational Animations and I also have my own AI Safety Channel. Link in the description. Despite researchers having a few proposed techniques that look promising, there's still one problem. Our goal is to use these techniques on AIs, which might be more capable than every single human expert. Can we really trust that the same techniques we use to train somewhat aligned book summarizers and chatbots today will just keep working on all models in the future? Ideally, we need some way to start testing whether our techniques will continue to work when AIs will surpass humans in every domain. We can test scalable oversight techniques right now. With this problem in mind, in 2021, Ajeya Cotra wrote a blog post titled "The Case for Aligning Narrowly Superhuman Models." She outlined the idea of "sandwiching": a way of evaluating scalable oversight techniques relative to superhuman AIs even before such AIs are built. She noticed that compared to average humans, current AIs show superhuman abilities on some narrowly defined tasks. In particular, recently we've started building models that are getting better than some humans on abstract, multi-dimensional, real world tasks that she calls "fuzzy tasks." These are tasks that don't have easily generated training signals or are hard to evaluate, like picking good stocks or giving good medical advice. Importantly, these models are not "superhuman" at these tasks in the same way that AlphaGo is "superhuman" at playing Go. AlphaGo plays Go better than any human, while GPT-4 is only capable of giving better advice or writing better stories than some humans. We still have domain specific experts that can outperform GPT-4. Importantly, the fact that such narrowly superhuman models exist means that we can take advantage of them to get some practice in dealing with superhuman AI that might help with more general systems. Ajeya Cotra suggested carrying out case studies on the various proposed scalable oversight techniques using one of these models, alongside humans that are less capable than the AI on some fuzzy task. This allows us to see how well different scalable oversight techniques handle situations where the AI model is more capable than the humans trying to align it. The aim here is to simulate the situation we expect to find ourselves in, in the future. So we have two groups of humans, experts and non-experts. The experts are standing in for ground truth. They have a good understanding of the relevant domain so they can tell us what the answers actually are. The non-experts are stand ins for a future version of humanity. They don't have a full understanding of the domain, but they do have various tools and techniques at their disposal. They need to use whatever tools they have to somehow oversee the outputs of an untrustworthy but superhuman system. Based on whether these oversight techniques succeed or fail in these case studies, we can start iteratively tweaking them, making them stronger over time. Using this process, we may be able to develop techniques that actually work when AIs will be better than the best humans at every task. Let's imagine sandwiching in practice. To get some intuition about how sandwiching might work, let's think about one of our current models, GPT-4. Imagine that you're sick and you want to ask for some advice. GPT-4 has seen huge amounts of information from all over the internet, including a lot of medical research. It seems plausible that GPT-4 could give better health advice than a random internet user on Reddit or Facebook. But all the language model is trying to do by default is to predict the next word as it would have appeared in the training data. It doesn't really "want" to give you good medical advice. In fact, text that constitutes bad advice might be considered more likely if it's seen more frequently online. You can explicitly ask it to give you good advice, but it's not really trying to follow instructions. It's trying to predict what comes next and in the training data, a request for "good advice" is still often followed by bad advice. So the AI's objective of generating probable text is misaligned with the user's objective, specifically for accurate and helpful information. To try to fix this problem, suppose that we gather a group of people who don't know much about medicine. We can arm them with one of our proposed scalable oversight techniques, and see how far they get in trying to make GPT-4 give useful medical advice. We can then compare performance to control groups like GPT-4 acting alone or unaugmented humans without any oversight techniques. The key point here is to measure how well these non-experts can get GPT-4 to give accurate advice without getting any actual medical experts involved. Basically, we want to see how well non-experts can independently take an unaligned model that is more capable than them and make it behave in a more aligned manner. We can then compare the result with the expert's judgment, or with the model aligned by the experts. In this case, we want the models to move in the direction of the accuracy and helpfulness that a qualified doctor might be able to provide. Any progress in this direction indicates some amount of success. It would constitute one data point suggesting that the oversight technique we tested might continue to work on strongly superhuman AIs. And now a real sandwiching experiment. In a 2022 paper titled "Measuring Progress on Scalable Oversight for Large Language Models," a group of researchers tested out how well sandwiching might work. In their experiment, they hired non-experts to try to solve two different types of tasks. One task was answering multiple-choice questions on the Massive Multitask Language Understanding benchmark or the MMLU for short. This has hundreds of questions, ranging from college level physics to abstract questions on morality. The other task was answering questions about short stories on a five minute deadline, when such stories would ordinarily take 15 to 30 minutes to read in full. Participants were asked to answer the questions and to rate how confident they were in the correctness of each answer with the goal of reaching the highest level of accuracy on the questions from the two datasets. The researchers made some simplifying assumptions for the experimental setup. The oversight technique was a simple back and forth text conversation with a language model to help augment the knowledge of the participants. And instead of human answers, they used the correct answers from the two tests. Despite the simplifications, the core of the sandwiching paradigm was preserved. When measured individually, the model outperformed unassisted humans on both tasks by a large margin. This means the basic condition for testing sandwiching, the model being better than the non-experts, was fulfilled. Over the course of the experiment, the participants got pretty good at using the chatbot to probe for facts. They also learned to break down complicated questions into simpler parts, which helped them understand the chatbot's answers better. So the assisted humans got substantially better scores than either unassisted humans or the model on its own. They didn't manage to match expert level performance estimated in other studies, though. There are also some problematic elements that the researchers noticed during the course of the experiment. For example, the chatbot sometimes agreed too easily with whatever the participants said, rather than correcting them when needed. Also, since the participants had limited domain knowledge and couldn't use external sources to fact-check, they accepted false claims as long as the chatbot sounded confident and the answer seemed plausible. This made them give highly confident judgements that turned out to be wrong. Despite the problems, simplifications, and the relatively unrealistic setting of multiple-choice questions, the participants did manage to move the behavior of the model in the direction that we would want. So the researchers effectively demonstrated sandwiching with this experimental design. This paper built a baseline demonstration that future experiments could refine, for example, by letting people fine tune the model, or implementing techniques such as debate and recursive reward modeling, or even letting the participant have access to interpretability techniques to better evaluate what the model says by looking at its internals. And sandwiching is not the only contender for ways to evaluate scalable oversight techniques. There are also other proposals being explored to ensure that any oversight processes we come up with today are going to scale to dangerously powerful future models, such as meta level adversarial evaluations. These techniques, and others like them, are helping to build an empirical science of AI safety. If we combine the development of new oversight techniques with improved evaluations of their efficacy, then we can start building up some justified confidence in our ability to oversee superhuman AIs in the future. If you'd like to get involved in the science of AI safety, AIsafety.com is a hub for navigating the AI existential safety ecosystem with many resources, such as the communities list with in-person and online groups, the events and training programs tracker, lists of funding opportunities, volunteer projects, a donation guide, and lots more. It also has a curated video playlist on AI safety featuring many of our videos. Also, you can support Rational Animations on Patreon and by buying our merch. We have T-Shirts, hoodies, stickers, notebooks and much more. Just head to rationalanimations.com and click on the "shop" tab. There's also a link in the description.
Video 8.2: Optional video explaining the concept of sandwiching (Rational Animations, 2025).
Sandwiching is an experimental setup that leverages narrowly superhuman models to evaluate how well our oversight techniques work in controlled settings. The sandwiching setup contains three primary ‘layers’:
- Non-Expert Layer: At the bottom, we have non-experts. These are people who wouldn't be able to solve the given task independently. As an example, think of lay people being asked to pick stocks or give medical diagnoses. These are stand-ins for a future version of humanity. They need to use the scalable oversight techniques to guide the AI's behavior in a more aligned direction.
- AI Model: In the middle, we have the narrowly superhuman AI model that we want to test the oversight technique on. This model has some capabilities that we are unable to elicit. The non-expert should be able to “align” the model by getting it to display the kind of behavior we want using the given scalable oversight technique.
- Expert Layer: At the top, we have domain experts who evaluate the success of the alignment attempt. They serve as a benchmark, providing feedback without directly interacting with the non-experts or the AI during the process. We can also use ground truth labels from datasets as replacements for the expert layer instead of hiring actual experts.

Figure 8.23: (Bowman et al., 2022)
Empirical demonstration of sandwiching. Researchers conducted an experiment with non-experts tasked to answer multiple choice questions. The task was to answer multiple-choice questions from the Measuring Massive Multitask Language Understanding (MMLU) benchmark. This benchmark contains hundreds of questions on math, physics, morality, ethics and a range of topics. (Hendrycks et. al, 2021) They also had to answer questions based on short stories from QuALITY benchmark in 5 minutes. (Pang et al., 2021) The scalable oversight technique that they were given was a simple amplification using an AI assistant chatbot. The performance was measured by using the correct answer ground truth labels from the dataset.
When measured individually, the model outperformed unassisted humans on both tasks by a large margin. This means the basic condition for testing sandwiching was fulfilled - the model being better than the non-experts. Over the course of the experiment, the chatbot assisted humans got substantially better scores on the tasks than either the humans or the models did alone. (Bowman et al., 2022) Even though this was an extremely simplified design, the participants managed to move the behavior of the model in the direction that we would want, i.e. make the model more “aligned”. So the researchers managed to effectively demonstrate sandwiching as an experimental design. Building on this baseline, future experiments can evaluate the efficacy of more complex scalable oversight methods like re-training, fine-tuning, or debate.
The whole approach rests on strong models already holding skills that weak supervision can draw out rather than teach. Did you find that assumption convincing? Talk it over with the tutor.