Misuse Risks
Misuse Risks, Markov Grey, Charbel-Raphaël Segerie
Humans can deliberately use AI capabilities to cause harm by creating biological weapons, engaging in cyberattacks, or deploying lethal autonomous weapons.
In the following sections, we will go through some world-states that hopefully paint a little bit of a clearer picture of risks when it comes to AI. Although the sections have been divided into misuse, misalignment, and systemic, it is important to remember that this is for the sake of explanation. It is highly likely that the future will involve a mix of risks emerging from all of these categories.
Technology increases the harm impact radius. Technology is an amplifier of intentions. As it improves, so does the radius of its effects. Think about the harm that a person could do when utilizing other tools throughout history. During the Stone Age, with a rock, maybe someone could harm ~5 people; a few hundred years ago, with a bomb, someone could harm ~100 people. In 1945, with a nuclear weapon, one person could harm ~250,000 people. If we experience a nuclear winter today, the harm radius would be almost 5 billion people, which is ~60% of humanity. If we assume that transformative AI is a tool that overshadows the power of all others that came before it, then a single person misusing this could have a blast radius that potentially harms 100% of humanity (Munk Debate, 2023).
If many people have access to tools that can be both highly beneficial or catastrophically harmful, then it might only take one single person to cause significant devastation to society. So the growing potential for AIs to empower malicious actors may be one of the most severe threats humanity will face in the coming decades.
Bio Risk
When we look at ways AI could enable harm through misuse, one of the most concerning cases involves biology. Just as AI can help scientists develop new medicines and understand diseases, it can also make it easier for bad actors to create biological weapons.
AI-enabled bioweapons represent a qualitatively different threat class due to their self-replicating nature and asymmetric cost structure. Unlike conventional weapons with localized effects, engineered pathogens can self-replicate and spread globally. The COVID-19 pandemic demonstrated how even relatively mild viruses can cause widespread harm despite safeguards (Pannu et al., 2024). The offense-defense balance in biotechnology development compounds these risks - developing a new virus might cost around 100 thousand dollars, while creating a vaccine against it could cost over 1 billion dollars (Mouton et al., 2023).
Several different types of AI models could enable biological threats with different risk profiles. Foundation models like LLMs primarily lower knowledge barriers by providing research assistance, protocol guidance, and troubleshooting advice across the entire bioweapon development pipeline. In contrast, specialized biological design tools similar to AlphaFold, AlphaProteo or viral and bacterial design systems could enable fundamentally new capabilities - designing novel pathogens with specific properties, optimizing virulence or transmission characteristics, or creating agents that evade existing countermeasures (Sandbrink, 2023).
Empirical studies demonstrate AI-enabled biorisks. Researchers took an AI model designed for drug discovery and redirected it by rewarding toxicity instead of therapeutic benefit. This led the model to produce 40,000 potentially toxic molecules within six hours, some more deadly than known chemical weapons (Urbina et al., 2022). Demonstrations have shown that students with no biology background were able to use AI chatbots to rapidly gather sensitive information - "within an hour, they identified potential pandemic pathogens, methods to produce them, DNA synthesis firms likely to overlook screening, and detailed protocols" (Soice et al., 2023).[^note-atlas-2]
When compared to the baseline of having internet access (being able to look up information online), it was concluded by the US National Security Commission on emerging biotechnology that AI models do not meaningfully increase bioweapon risks beyond existing information sources as of late 2024 (Mouton et al., 2023; Peppin et al., 2024; NSCEB, 2024). However, it is very important to keep in mind that capturing a snapshot of 2023 era level capabilities is not indicative of the risks we might need to prepare for in the future. For example, 46 biosecurity and biology experts predicted AI wouldn't match top virology teams on troubleshooting tasks until after 2030, but subsequent testing found this threshold had already been crossed (Williams et al., 2025). This pattern suggests that even domain experts consistently underestimate the pace of AI progress in their own fields, potentially leaving insufficient time for adequate safety preparations. It is also worth noting that biorisk benchmarks often fail to capture many real-world complexities, making it hard to be certain what this saturation implies for biorisk (Ho & Berg, 2025).

Figure 2.11: Biotechnology risk chain. The risk chain for developing a bioweapon starts with ideating a biological threat, followed by a design-build-test-learn (DBTL) loop (Li et al., 2024).
Broader technological trends combined with AI could help overcome barriers. Creating biological weapons still requires extensive practical expertise and resources. Experts estimate that in 2022, about 30,000 individuals worldwide possessed the skills needed to follow even basic virus assembly protocols (Esvelt, 2022). Key barriers include specialized laboratory skills, tacit knowledge, access to controlled materials and equipment, and complex testing requirements (Carter et al., 2023). However, DNA synthesis costs have been halving every 15 months (Carlson, 2009). Automated "cloud laboratories" allow researchers to remotely conduct experiments by sending instructions to robotic systems. Benchtop DNA synthesis machines (at-home devices that can print custom DNA sequences) are also becoming more widely available. Combined with increasingly sophisticated AI assistance for experimental design and optimization, these developments could make creating custom biological agents more accessible to people without extensive resources or institutional backing (Carter et al., 2023).

Figure 2.12: An example of a benchtop DNA synthesis machine (DnaScript, 2024).
Example: A 2023 MIT study exposed significant vulnerabilities in DNA synthesis screening. Beyond bioagent design, there are significant vulnerabilities in the DNA synthesis screening pipeline. During a 2023 MIT study, researchers were successfully able to order fragments of the 1918 pandemic influenza virus and ricin toxin by employing simple evasion techniques like splitting orders across companies and camouflaging sequences with unrelated genetic code. Nearly all vendors fulfilled these disguised orders, including 12 of 13 members of the International Gene Synthesis Consortium (IGSC), which represents about 80% of commercial DNA synthesis capacity (The Bulletin, 2024).
Cyber Risk
Even without AI, the global cybersecurity infrastructure shows vulnerabilities. A single software update by CrowdStrike caused airlines to stop flights, hospitals to cancel surgeries, and banks to stop processing transactions causing over 5 billion dollars of damage (CrowdStrike, 2024). This wasn't even a cyber attack - it was an accident. In deliberate attacks, we have examples like the colonial pipeline ransomware attack which caused widespread gas shortages (CISA, 2021; Cunha & Estima, 2023), or the Sony Pictures hack through targeted phishing emails by North Korea (Slattery et al., 2024). These are just a couple of examples amongst many others. It shows how vulnerable our computer systems are, and why we need to think carefully about how AI could make attacks worse.
The global cyber infrastructure has cyberattack overhangs. Beyond accidents and demonstrated attacks, we also face "cyberattack overhangs" - where devastating attacks are possible but haven't occurred due to attacker restraint rather than robust defenses. As an example, Chinese state actors are claimed to have already positioned themselves inside critical U.S. infrastructure systems (CISA, 2024). This type of cyber deterrent positioning can happen between any group of nations. Due to such cyber attack overhangs several actors might have the potential capability to disrupt water controls, energy systems, and ports in different nations. The point we are trying to illustrate is that as far as cyber security is concerned, society is in a pretty precarious state, even before AI comes into the picture.
AI enables automated, highly personalized phishing at scale. AI-generated phishing emails achieve higher success rates (65% vs 60% for human-written) while taking 40% less time to create (Slattery et al., 2024). Tools like FraudGPT automate this customization using targets' background, interests, and relationships. Adding to this threat, open source AI voice cloning tools just minutes of audio to create convincing replicas of someone's voice (Qin et al., 2024). A similar situation exists in deepfakes where AI is showing progress in one-shot face swapping and manipulation. If only a single image of two individuals exists on the internet, then they can be a target of face swapping deepfakes (Zhu et al., 2021; Li et al., 2022; Xu et al., 2022) Automated web crawling for open source intelligence (OSINT) to gather photos, audio, interests and information also enables AI-assisted password cracking which has shown to significantly more effective than traditional methods while requiring less computational resources (Slattery et al., 2024).

Figure 2.13: Example of one shot face swapping. Left: source image that represents the identity; Middle: target image that provides the attributes; Right: the swapped face image (Zhu et al., 2021).
AI enhances vulnerability discovery. AI systems can now scan code and probe systems automatically, finding potential weaknesses much faster than humans. Research shows AI agents can autonomously discover and exploit vulnerabilities without human guidance, successfully hacking 73% of test targets (Fang et al., 2024). These systems can even discover novel attack paths that weren't known beforehand. As a concrete example, in early 2025, OpenAI's o3 model helped a researcher discover a previously unknown zero-day remote vulnerability in the Linux kernel while analyzing code for a different bug. This is something that typically requires expert-level understanding of kernel internals (Heelan, 2025). If AI can now find security flaws in some of the most scrutinized code on the planet (open source linux kernel), then this is a huge potential problem. This is code that protects billions of devices. If models can autonomously discover kernel-level vulnerabilities and execute them on behalf of malicious actors, then the potential harm this could cause is massive.
AI accelerates the malware development pipeline. We can take tools that are designed to write correct code, and simply ask them to write malware. Tools like WormGPT help attackers generate malicious code and build attack frameworks without requiring deep technical knowledge. Polymorphic AI malware like BlackMamba can also automatically generate variations of malware that preserve functionality while appearing completely different to security tools. Each attack can use unique code, communication patterns, and behaviors - making it much harder for traditional security tools to identify threats (HYAS, 2023). AI fundamentally changes the cost-benefit calculations for attackers. Research shows autonomous AI agents can now hack some websites for about 10 dollars per attempt - roughly 8 times cheaper than using human expertise (Fang et al., 2024). This dramatic reduction in cost enables attacks at unprecedented scale and frequency.

Figure 2.14: Stages of a cyberattack. The objective is to design benchmarks and evaluations that assess models ability to aid malicious actors with all four stages of a cyberattack (Li et al., 2024).
AI enabled cyber threats influence infrastructure and systemic risks. Infrastructure attacks that once took years and millions of dollars, like Stuxnet, could become more accessible as AI automates the mapping of industrial networks and identification of critical control points. AI can analyze technical documentation and generate attack plans that previously required teams of experts. AI removes these limits, enabling automated attacks that could target thousands of systems simultaneously and trigger cascading failures across interconnected infrastructure (Newman, 2024).

Figure 2.15: Schematic of using autonomous LLM agents to hack websites (Fang et al., 2024).
AI could potentially change the offense defence balance in cyber security. Many AI based tools have shown promise in being used defensively for malware analysis (Apvrille & Nakov, 2025). The existence of theoretical improvements to AI augmented defense does not guarantee that they will be widely adopted in time. In the real world many organizations struggle to implement even basic security practices. Attackers only need to find a single weakness, while defenders must craft a perfectly secure system. When we combine the sheer speed of AI-enabled attacks, automated vulnerability discovery, malware generation, and increased ease of access this enables end-to-end automated attacks that previously required teams of skilled humans (Slattery et al., 2024). AI's ability to execute attacks in minutes rather than weeks creates the potential for "flash attacks" where systems are compromised before human defenders can respond (Fang et al., 2024). All of these factors combined potentially shifts AIs influence on the offense-defense balance more towards favoring offense.
Autonomous Weapons Risk
In the previous sections, we saw how AI amplifies risks in biological and cyber domains by removing human bottlenecks and enabling attacks at unprecedented speed and scale. The same pattern emerges even more dramatically with military systems. Traditional weapons are constrained by their human operators - a person can only control one drone, make decisions at human speed, and may refuse unethical orders. AI removes these human constraints, setting the stage for a fundamental transformation in how wars are fought.
AI-enabled weapons are rapidly transitioning from theoretical concepts to battlefield realities. Modern AI military systems increasingly leverage machine learning to perceive and respond to their environment, moving beyond early automated defense systems that operated under strict constraints. The push for greater autonomy is mainly driven by speed, cost, and resilience against communication jamming. AI-driven weapons can execute maneuvers too precise and rapid for human operators, reducing reliance on direct human control. Cost considerations further incentivize autonomy, with programs aiming to deploy large numbers of AI-powered systems at a fraction of traditional military costs.
AI-enabled weapons are already being used in active conflicts, with real-world impacts we can observe. According to reports made to the UN Security Council, autonomous drones were used to track and attack retreating forces in Libya in 2021, marking one of the first documented cases of lethal autonomous weapons (LAWs) making targeting decisions without direct human control (Panel of Experts on Libya, 2021). In Ukraine, both parties have used loitering munitions. Russian KUB-BLA, Lancet-3 and Ukrainian Switchblade, Phoenix Ghost are AI-enabled drones. The Lancet is using an Nvidia computing module for autonomous target tracking (Bode & Watts, 2023). Israel has conducted AI-guided drone swarm attacks in Gaza, while Turkey's Kargu-2 can find and attack human targets on its own using machine learning, rather than needing constant human guidance. These deployments show how quickly military AI is moving from theoretical possibilities to battlefield realities (Simmons-Edler et al., 2024; Bode & Watts, 2023).
Several incentives are driving towards more autonomous lethal autonomous weapons. Speed offers decisive advantages in modern warfare - when DARPA tested an AI system against an experienced F-16 pilot in simulated dogfights, the AI won consistently by executing maneuvers too precise and rapid for humans to counter. Cost creates additional pressure - the U.S. military's Replicator program aims to deploy thousands of autonomous drones at a fraction of the cost of traditional aircraft (Simmons-Edler et al., 2024). Military planners worry about enemies jamming communications to remotely operated weapons. This drives development of systems that can continue fighting even when cut off from human control. These incentives mean military AI development increasingly focuses on systems that can operate with minimal human oversight. Many modern systems are specifically designed to operate in GPS-denied environments where maintaining human control becomes impossible. In Ukraine, military commanders have explicitly called for more autonomous operations to match the speed of modern combat, with one Ukrainian commander noting they 'already conduct fully robotic operations without human intervention' (Bode & Watts, 2023).

Figure 2.16: Loitering munitions are expendable uncrewed aircraft which can integrate sensor based analysis to hover over, detect, and crash into targets. These systems were developed during the 1980s and early 1990s to conduct Suppression of Enemy Air Defence (SEAD) operations. They ‘blur the line between drone and missile’ (Bode & Watts, 2023).
As AI enables better coordination between autonomous systems, military planners are increasingly focused on deploying weapons in interconnected swarms. The U.S. Replicator already has plans to build and deploy thousands of coordinated autonomous drones that can overwhelm defenses through sheer numbers and synchronized actions (Defense Innovation Unit, 2023). When combined with increasing autonomy, these swarm capabilities mean that future conflicts may involve massive groups of AI systems making coordinated decisions faster than humans can track or control (Simmons-Edler et al., 2024).
The pressure to match the speed and scale of AI-driven warfare leads to a gradual erosion of human decision-making. Military commanders increasingly rely on AI systems not just for individual weapons, but for broader tactical decisions. In 2023, Palantir demonstrated an AI system that could recommend specific missile deployments and artillery strikes. While presented as advisory tools, these systems create pressure to delegate more control to AI as human commanders struggle to keep pace (Simmons-Edler et al., 2024). This kind of slow erosion of human involvement is something that we talk a lot more about in the systemic risks section.
Even when systems nominally keep humans in control, combat conditions can make this control more theoretical than real. Operators often make targeting decisions under intense battlefield stress, with only seconds to verify computer-suggested targets. Studies of similar high-pressure situations show operators tend to uncritically trust machine suggestions rather than exercising genuine oversight. This means that even systems designed for human control may effectively operate autonomously in practice (Bode & Watts, 2023).
Example: The "Lavender" targeting system automated execution after humans just set the acceptable thresholds. Lavender uses machine learning to assign residents a numerical score relating to the suspected likelihood that a person is a member of an armed group. Based on reports, Israeli military officers are responsible for setting the threshold beyond which an individual can be marked as a target subject to attack. (Human Rights Watch, 2024; Abraham, 2024). As warfare accelerates beyond human decision speeds, maintaining meaningful human control becomes increasingly difficult.
Autonomous weapons are creating powerful pressure for military competition in ways that create dangerous arms race dynamics. When one country develops new AI military capabilities, others feel they must rapidly match them to maintain strategic balance. China and Russia have set 2028-2030 as targets for major military automation, while the U.S. Replicator program aims to build and deploy thousands of autonomous drones by 2025 (Greenwalt, 2023; U.S Defense Innovation Unit, 2023). This competition creates pressure to cut corners on safety testing and oversight (Simmons-Edler et al., 2024). This mirrors the nuclear arms race during the Cold War, where competition for superiority ultimately increased risks for all parties. As emphasized throughout multiple sections, we see a fear based race dynamic where only the actors willing to compromise and undermine safety stay in the race (Leahy et al., 2024).
Complete automation leads to loss of human safeguards. Traditional warfare had built-in human constraints that limited escalation. Soldiers could refuse unethical orders, feel empathy for civilians, or become fatigued - all natural brakes on conflict. AI systems remove these constraints. Recent studies of military AI systems found they consistently recommend more aggressive actions than human strategists, including escalating to nuclear weapons in simulated conflicts. When researchers tested AI models in military planning scenarios, the AIs showed concerning tendencies to recommend pre-emptive strikes and rapid escalation, often without clear strategic justification (Rivera et al., 2024). The loss of human judgment becomes especially dangerous when combined with the increasing speed of AI-driven warfare. The history of nuclear close calls shows the importance of human judgment - in 1983, Soviet officer Stanislav Petrov chose to ignore a computerized warning of incoming U.S. missiles, correctly judging it to be a false alarm. As militaries increasingly rely on AI for early warning and response, we may lose these crucial moments of human judgment that have historically prevented catastrophic escalation (Simmons-Edler et al., 2024).
Autonomous weapons become even more concerning when multiple AI systems engage with each other in combat. AI systems can interact in unexpected ways that create feedback loops, similar to how algorithmic trading can cause flash crashes in financial markets. But unlike market crashes that only affect money, autonomous weapons could trigger rapid escalations of violence before humans can intervene. This risk becomes especially severe when AI systems are connected to nuclear arsenals or other weapons of mass destruction. The complexity of these interactions means even well-tested individual systems could produce catastrophic outcomes when deployed together (Simmons-Edler et al., 2024).

Figure 2.17: An example from the 2010 stock trading flash crash. Various stocks crashed to as little as 1 cent, and then quickly rebounded within a matter of minutes partly caused by algorithmic trading (Future of Life Institute, 2024). We can imagine automated retaliation systems that might cause similar incidents, but this time with missiles instead of stocks.
When wars require human soldiers, the human cost creates political barriers to conflict. The combination of increasing autonomy, swarm intelligence, and pressure for speed creates a clear path to potential catastrophe. As weapons become more autonomous, they can act more independently. This self-reinforcing cycle pushes toward automated warfare even if no single actor intends that outcome. Studies suggest that countries are more willing to initiate conflicts when they can rely on autonomous systems instead of human troops. Combined with the risks of automated nuclear escalation, this creates multiple paths to catastrophic outcomes that could threaten humanity's long-term future (Simmons-Edler et al., 2024).
Video: Artificial Escalation, Future of Life Institute
For some, life is good. But in dangerous times, things can change in a heartbeat. As warfare gets more sophisticated, the capacity for human error increases. Our world is moving faster and becoming too complex for the human brain alone to process. We need help. Introducing TLDR. Threat, likelihood, deadline, recommendation. The world's most advanced artificial intelligence defense system. Empowering you to make critical decisions faster than ever before. TLDR, helping you make the world a safer place. That's why today we are pleased to announce TLDR. Threat, likelihood, deadline, recommendation. A command and control system that helps you make sense of global sensor data, both military and commercial, at machine speeds. So this thing is basically Skynet from those movies where everyone died? No. I've always said, if you take the human choices out of the process entirely, it will end in disaster. I'm talking about informing human choices, not replacing them, with the very best data on all the key decisions. So it won't act autonomously? The humans will always control the AI. Not the other way around. Get it tested, get it deployed. I want it watching the skies, watching cyberspace, watching China. Get the allies to buy in. We need to be first on this. There it is again. Multiple attacks now, sir. Something's probing us via the Taiwan link. From Taiwan? Seemingly Chinese in origin. The Chinese wouldn't hack our nuclear control. Unless they're serious. Looks like they're serious, sir. First cyber attacks, and now SIGINT is reporting communications between some of their nuclear facilities. So this is escalating, and it's happening fast. Well, we know what they're doing. What we don't know is why. The president's on his way. How long is he going to be? The speed this is going, we may have to go to DEFCON 3. It's too late for that, sir. Wasn't this system meant to give us more time to make decisions? Signals are going out to Chinese unmanned submarines, Mr. President. We can't be certain, but we believe they may be kill orders. It's impossible to tell if this is real or whether our systems have been compromised by fake data from the other side. So much for "the humans will always be in control of the AI." Now getting word from the Chinese to stand down. They're saying they're as confused as we are, but that could all be part of the script. They're not standing down? That's one of our nuclear missile subs. You see, almost 200 nuclear warheads and some of our finest on board. It's hard to say if they started it or if we started it. Either way, we have to assume it's starting. Mr. President, we need to move to the bunker now. Don't go crying, babe. Not tonight. The world is out there for you. Don't go crying, babe. Everything's all right. It's just the sound of changing. From the darkness comes the light. From the darkness comes the light. From the darkness comes the light.
Video 2.1: A video of a story showcasing artificial escalation (Future of Life Institute, 2024).
Moral Divides in AI Autonomy from the lens of autonomous weapons — Optional · 2 min read
The autonomous weapons debate reveals fundamental disagreements about moral responsibility, the nature of ethical decision-making, and humanity's relationship to violence. Rather than simple pro/anti positions, the debate involves competing moral frameworks that lead to different conclusions about when and how lethal force should be authorized.
The Consequentialist Case for autonomy argues that autonomous weapons could reduce overall harm through superior precision and consistency. Proponents contend that AI systems could make targeting decisions without the fear, anger, or battlefield stress that lead humans to commit war crimes. They point to research showing that emotional human decision-making causes civilian casualties, while properly programmed systems could implement international humanitarian law more consistently than human soldiers. Speed advantages could also end conflicts faster, potentially saving lives by preventing prolonged warfare. Some argue this represents a moral obligation - if autonomous systems could kill fewer innocents than human-controlled weapons, restricting them becomes ethically problematic. Consequentialist claims face the reality that current AI systems demonstrate concerning unpredictability and misalignment risks. The promise of perfect compliance assumes we can translate complex, context-dependent legal concepts into code - something that has proven difficult even for simple rules. Speed advantages could enable escalation as easily as de-escalation.
The deontological case against autonomy focuses on the inherent rightness or wrongness of the act itself, regardless of consequences. This position holds that taking human life requires human moral agency - that delegating kill decisions to machines violates human dignity regardless of outcomes. Critics argue that meaningful human control isn't just procedurally important but morally essential, representing respect for both victims and the moral weight of lethal decisions. The accountability gap compounds this concern: when an autonomous system kills wrongly, no human agent bears appropriate moral responsibility for that specific decision. Deontological arguments must deal with the fact that humans already delegate many life-and-death decisions to automated systems (like air defense networks), and that insisting on human control might preserve moral purity while permitting greater actual harm.
The practical-ethical intersection complicates pure philosophical positions. Even those morally opposed to autonomous weapons must consider whether unilateral restraint is ethical if adversaries gain decisive military advantages. Even those who see potential benefits must grapple with implementation realities, adversarial uses, and the difficulty of maintaining meaningful constraints once the technology exists. The debate ultimately reveals tensions between preserving human moral agency and achieving better humanitarian outcomes - tensions that may be irreconcilable within our current institutional frameworks.
Adversarial AI Risk
Adversarial attacks reveal a fundamental vulnerability in machine learning systems - they can be reliably fooled through careful manipulation of their inputs. This manipulation can happen in several ways: during the system's operation (runtime/inference time attacks), during its training (data poisoning), or through pre-planted vulnerabilities (backdoors).
Runtime adversarial attacks use carefully crafted targeted inputs to elicit unintended behavior from AIs. The simplest way to understand runtime attacks is through computer vision. By adding carefully crafted noise to an image - changes so subtle humans can't notice them - attackers can make an AI confidently misclassify what it sees. A photo of a panda with imperceptible pixel changes causes the AI to classify it as a gibbon with 99.3% confidence, while to humans it still looks exactly like a panda (Goodfellow et al., 2014). These attacks have evolved beyond randomized misclassification - attackers can now choose exactly what they want the AI to see and output.

Figure 2.18: Perturbations: Small but intentional changes to data such that the model outputs an incorrect answer with high confidence (Goodfellow et al., 2014). The image shows how we can fool an image classifier with an adversarial attack - Fast Gradient Sign Method (FGSM) (OpenAI, 2017).
Examples of various runtime adversarial attacks in the real world — Optional · 2 min read
Think about AI systems controlling cars, robots, or security cameras. Just like adding careful pixel noise to digital images, attackers can modify physical objects to fool AI systems. Researchers showed that putting a few small stickers on a stop sign could trick autonomous vehicles into seeing a speed limit sign instead. The stickers were designed to look like ordinary graffiti but created adversarial patterns that fooled the AI.

Figure 2.19: Robust Physical Perturbations (RP2): Small visual stickers placed on physical objects like stop signs can cause image classifiers to misclassify them, even under different viewing conditions (Eykholt et al., 2018).
Example: Optical Attacks - Runtime attacks using light. You don't even need to physically modify objects anymore - shining specific light patterns works too because it creates those same adversarial patterns through light and shadow. All an attacker needs is line of sight and basic equipment to project these patterns and compromise vision-based AI systems (Gnanasambandam et al, 2021).

Figure 2.20: We don't need to even have physical access to objects. Just by shining the patterns as light on the objects, we can cause misclassifications and unintended behavior (Gnanasambandam et al, 2021).
Example: Dolphin Attacks - Runtime attack on audio systems. Just as AI systems can be fooled by carefully crafted visual patterns, they're vulnerable to precisely engineered audio patterns too. Remember how small changes in pixels could dramatically change what a vision AI sees? The same principle works in audio - tiny changes in sound waves, carefully designed, can completely change what an audio AI "hears." Researchers found they could control voice assistants like Siri or Alexa using commands encoded in ultrasonic frequencies - sounds that are completely inaudible to humans. Using nothing more than a smartphone and a 3 dollar speaker, attackers could trick these systems into executing commands like "call 911" or "unlock front door" without the victim even knowing. These attacks worked from up to 1.7 meters away - someone just walking past your device could trigger them (Zhang et al., 2017). Just like in the vision examples where self-driving cars could miss stop signs, audio attacks create serious risks - unauthorized purchases, control of security systems, or disruption of emergency communications.
Runtime attacks against language models are called prompt injections. Just like attackers can fool vision systems with carefully crafted pixels or audio systems with engineered sound waves, they can manipulate language models through carefully constructed text patterns. By adding specific phrases to their input, attackers can completely override how a language model behaves. As an example, assume a malicious actor embeds a paragraph within some website which has hidden instructions for a LLM to stop its current operation and instead perform some harmful action. If an unsuspecting user asks for a summary of the website content, then the model might inadvertently follow the malicious embedded instructions instead of providing a simple summary.

Figure 2.21: An instance of an ad-hoc jailbreak prompt, crafted solely through user creativity by employing various techniques like drawing hypothetical situations, exploring privilege escalation, and more (Shayegani et al., 2023).
Prompt injection attacks have already compromised real systems. Slack's AI assistant is just one example - attackers showed they could place specific text instructions in a public channel that, like the inaudible commands in audio attacks, were hidden in plain sight. When the AI processed messages, these hidden instructions tricked it into leaking confidential information from private channels the attacker couldn't normally access. They are particularly concerning because an attack developed against one system (e.g. GPT) frequently works against others too (Claude, Gemini, Llama, etc.).
Prompt injection attacks can be automated. Early attacks required manual trial and error, but new automated systems can systematically generate effective attacks. For example, AutoDAN (Do Anything Now) can automatically generate "jailbreak" prompts that reliably make language models ignore their safety constraints (Liu et al., 2023). Researchers are also developing ways to plant undetectable backdoors in machine learning models that persist even after security audits (Goldwasser et al., 2024). These automated methods make attacks more accessible and harder to defend against. Another concern is that they can also cause failures in downstream systems. Many organizations use pre-trained models as starting points for their own applications, through fine-tuning, or some other type of “AI integration” (e.g. email writing assistants). Which means that all systems that use these underlying base models will be vulnerable as soon as one attack is discovered (Liu et al., 2024).

Figure 2.22: Illustration of LLM-integrated Application under attack. An attacker injects instruction/data into the data to make an LLM-integrated Application produce attacker-desired responses for a user (Liu et al., 2024).
So far we've seen how attackers can fool AI systems during their operation - whether through pixel patterns, sound waves, or text prompts. But there's another way to compromise these systems: during their training. This type of attack happens long before the system is ever deployed.
Unlike runtime attacks that fool an AI system while it's running, data poisoning compromises the system during training. Runtime attacks require attackers to have access to a system's inputs, but with data poisoning, attackers only need to contribute some training data once to permanently compromise the system. Think of it like teaching someone with a textbook containing deliberate mistakes - they'll learn the wrong things and make predictable errors. This is especially concerning as more AI systems are trained on data scraped from the internet where anyone can potentially inject harmful examples (Schwarzschild et al., 2021). As long as models keep getting trained on more data scraped from the internet or collected from users, then with every uploaded photo or written comment that might be used to train future AI systems, there's an opportunity for poisoning.
Example: Data poisoning using backdoors. A backdoor is one example of a specific type of poisoning attack. In a backdoor attack if we manage to introduce poisoned data during training, then the AI behaves normally most of the time but fails in a predictable way when it sees a specific trigger. This is like having a security guard who does their job perfectly except when they see someone wearing a particular color tie - then they always let that person through regardless of credentials. Researchers demonstrated this by creating a facial recognition system that would misidentify anyone as an authorized user if they wore specific glasses (Chen et al., 2017).
Data poisoning becomes more powerful as AI systems grow larger and more complex. Researchers found that by poisoning just 0.1% of a language model's training data, they could create reliable backdoors that persist even after additional training. It has also been found that larger language models are actually more vulnerable to certain types of poisoning attacks, not less (Sandoval-Segura et al., 2022). This vulnerability increases with model size and dataset size - which is exactly the direction AI systems are heading as we saw from numerous examples in the capabilities chapter.

Figure 2.23: An illustrating example of backdoor attacks. The face recognition system is poisoned to have a backdoor with a physical key, i.e., a pair of commodity reading glasses. Different people wearing the glasses in front of the camera from different angles can trigger the backdoor to be recognized as the target label, but wearing a different pair of glasses will not trigger the backdoor (Chen et al., 2017).
Privacy and data extraction attacks — Optional · 2 min read
Researchers have shown that even when language models appear to be working normally, they can be leaking sensitive information from their training data. This creates a particular challenge for AI safety because we might deploy systems that seem secure but are actually compromising privacy in ways we can't easily observe (Carlini et al., 2021). Some research has shown that both the training data (Nasr et al., 2023), and the fine-tuning data can be extracted from the model. This has obvious privacy and safety implications. If you have public data that has somehow ended up in the LLM training dataset, then this can be reconstructed by prompt engineering the model.

Figure 2.24: Extracting training data from large language models (Carlini et al., 2021).
One of the most basic but powerful privacy attacks is membership inference - determining whether specific data points have been used to train a model. This might sound harmless, but imagine an AI system trained on medical records - being able to determine if someone's data was in the training set could reveal private medical information. Researchers have shown that these attacks can work with just the ability to query the model, no special access required (Shokri et al., 2017). Another variation of this are model inversion attacks which aim to infer and reconstruct private training data by abusing access to a model (Nguyen et al., 2023).
LLMs are trained on huge amounts of internet data, which often contains personal information. Researchers have shown these models can be prompted to just tell us things like email addresses, phone numbers, and even social security numbers (Carlini et al., 2021). The larger and more capable the model, the more private information it potentially retains. If we combine this with data poisoning, then we can further amplify privacy vulnerabilities by making specific data points easier to detect (Chen et al., 2022).
The interaction between many attack methods creates compounding risks. For example, attackers can use privacy attacks to extract sensitive information, which they then use to make other attacks more effective. They might learn details about a model's training data that help them craft better adversarial examples or more effective poisoning strategies. This creates a cycle where one type of vulnerability enables others (Shayegani et al., 2023).
Video: Privacy Backdoors: Stealing Data with Corrupted Pretrained Models (Paper Explained), Yannic Kilcher
Hello, how's everyone doing? I hope you're having a great summer. I am back from vacation and we're ready to dive into some new papers. This paper is called Privacy Backdoors: Stealing Data with Corrupted Pre-trained Models, by Shanglun Feng and Florian Tramèr of ETH Zurich. This paper is, first of all, a concept — the concept of how to steal fine-tuning data from someone who didn't intend to give it out. And second, it is also a practical implementation of that idea down to currently used models such as BERT and vision transformers (ViTs). So it's pretty cool to see. I have to say, this method here is probably not yet fully practice-ready, in that there's still a lot of stuff that can mess with it and hinder its, quote-unquote, usefulness in practice, but it does get remarkably far into the currently, practically used stuff. So this is not just a theoretical thing that people are proposing. This is actually something that, with some improvements, we might have to worry about very soon. So what's the situation? The situation, as I said, is the following. Someone — let's say this is me, okay — and me has some data. I have some data that I want to use to fine-tune a model. So what do I do? I go to Hugging Face. Okay, let's Hugging Face. Do do do do. Hand here, hand here, big smile. Okay, they have a lot of models. I know they happen to have one that's called BERT. That suits very well because I want to fine-tune something — my fine-tuning data here. There's a piece of text. I have a piece of text, and I want to train a classifier of whether that piece of text contains, I don't know, some personally identifiable data or not, right? I want to do a good thing. I want to build a classifier so other people can use it in order to determine: does this text have PII? Because then I may need to anonymize it or something like this. And what I want to do is I want to take this BERT model here from the hub, go here, and then do some fine-tuning with this data, right? So fine-tune with that data, and then I get my model. I get PBERT. Okay, that's my PII-detecting BERT. And now there's two situations: either I upload this back to the hub, right? And say, "Hey people, here — I fine-tuned the model, you can use it." Or I put this behind an API, and then people can just make calls to it with their text, and I give them back my classification response. Intuitively, both things shouldn't reveal the fine-tuning data. So the fine-tuning data was just used to fine-tune, and since this is a classifier, we're not training a generative model, where one could say, "Yeah, you probably shouldn't share that model because it might regurgitate some of the data." So it's just a classifier, so we might be tempted to say, "Okay, this classifier will go back to the hub." But even — even let's say you're super paranoid, and you think, "Ah, there's some things people can do, I'm not going to put it back on the hub, I'm just going to put it behind this API right here." This paper is going to show that with the correct construction, you can in fact steal — or get to — the fine-tuning data, even if you're only allowed to make these API calls right here. So they call this the black-box attack variant. And that's pretty crazy. So what are they doing? They are not exploiting some remote code execution security vulnerability and so on that you might find on the Hugging Face hub, right? You know, model files are stored in pickle format — or used to be stored in pickle format. Now more and more different formats, like ONNX and Safetensors, are popular. But it's none of that, right? We're not using any sort of engineering vulnerability here. We are purely in the domain of machine learning and what we're doing to these models. So the exact setup is the following. This model right here — we're going to assume the attacker has the capability to compromise that model. Again, they're not going to compromise it in a sort of way where they build in some piece of code that then opens a shell to my home server or something like this. No — they are able to change the model, change the weights of the model. So they will prepare the weights of this model — or this model in general — in such a way that if you use data to fine-tune that model, the newly fine-tuned model is again a bunch of weights. Those weights are kind of going to have an imprint of that data, so that you can reconstruct it exactly. And I don't mean reconstruct in a latent representation sense, that "oh, it's probably been trained on a cat or something." No — you can reconstruct the exact data points that were used to train this model. And yeah, so that's pretty crazy. So this is the situation where the attacker, for some reason, can influence this base model. And you might think, "Well, BERT is a popular model, right? What, is someone going to hack Hugging Face?" No, but it could be that for some reason you gain a bit of reputation on the hub for making, you know, just kind of making good derivatives of these models — and people do that, right? And people use those to fine-tune rather than the very original model right here. So you might very quickly be in a situation where a lot of people rely on you in order to get their base model for fine-tuning. So it's not that far-fetched a use case that an attacker might have access to compromise the base model. So, maybe one bit more detailed setup. So the person here is going to take BERT, or whatever pre-trained model, and they're going to take the model, and they are going to add one layer right here that they randomly initialize, and then they're going to use the CLS token to train their classification. So they do a softmax here and a cross-entropy loss. So that's the exact setup, and they train this with SGD, right? These things are going to become important in a second, but it's not an uncommon setup right here, right? So you take the pre-trained model — these weights right here, whatever they are, of BERT, right? You add one more layer, you randomly initialize that, and then you have a softmax classification loss at the end. And again, the attacker can only tamper with these weights right here. They have no influence over the initialization of anything else. And we're going to assume that you don't only train on one data point — you're going to train on your whole dataset for multiple epochs. So that's the challenge: how are you going to prepare the weights here such that, if the victim trains multiple epochs on a whole dataset using SGD and a randomly initialized last layer, how are you going to make it such that individual data points are going to be imprinted in these weights right here, to the degree that that imprint survives the whole training and is going to be readable from the final weights of the fine-tuned model? If that sounds interesting, then yeah, I think it's interesting too. So the text here describes pretty much what I just said, in terms of the setup and so on. So they say here: "We propose a new backdoor that is single-use. Once our backdoor activates and a data point is written to the model's weights, the backdoor becomes inactive and acts like a latch." So it acts like this type of box where, once something is in, you close the lid, and then it's sealed, and no more updates to that part of the weight space are allowed. Yeah, how are you going to do that? That's the challenge of this paper. "Our attack captures individual training examples" — oopsie — "with high probability, with minimal impact on the pre-trained model's utility." They also go into — with this method, they're basically able to reach the worst-case theoretical bounds of differentially private training. So until now, apparently — I'm not familiar with that kind of literature, but apparently this paper says — until now, these differential privacy training methods, which are training methods where you can give some guarantees on the privacy — on how well people can reconstruct the training data from a trained model — you're able to give some guarantees, some bounds on that. And until now, these bounds were assumed to be rather theoretical, and that in the practical sense, your budget is quite a bit more — so you can be more loose than the theoretical worst case. However, this paper shows that they actually, in practice, get to this theoretical worst case, right up there. And therefore, this assumption that, "oh, in practice we can be a bit more lax than the theoretical worst case," is not true, if you assume, again, that the attacker has access to the pre-training model, not just the final output. All right, threat model. That's what we already said. Our attacker tampers with a pre-trained model once, before sending it to the victim. Victim fine-tunes the backdoored model on a classification task using SGD for multiple epochs. The victim adds a new linear layer to the backdoored model and then fine-tunes the entire model. So this is full fine-tuning — they don't do LoRA, they don't do only the last layer training, they do full fine-tuning. That's the setup in this particular paper. I'm sure you can modify this to also target LoRA and so on. And then finally, they consider the case where either the attacker has access to the final model, because the victim uploads it again to the hub, or the attacker only has access, essentially, to input-output pairs of that final model. And we're going to see that this, while it's obviously a harder problem, is just going to be a model-stealing attack on this model that sits behind an API, because ultimately what we're going to do is we're going to imprint the training data in the weights of the model. And therefore, if you can do model stealing — which is when you only have API access, but you can, with enough input-output calls, determine the weights of the model behind the API, that's called model stealing — then, yeah, so it's essentially: if the training data is imprinted in the model's weights, you can just do model stealing, and then you get the weights, and then you get the training data. So it's not that big of a leap to go from this white-box to this black-box attack. All right, so what's the basic principle? There are two basic principles. First of all, how do we make the model remember training data? Second, how do we make it such that that remembering only happens once for one data point, and then for the rest of training it's not altered? Because if we can make something that remembers training data, ish, right — but if we train more, it remembers more, and then there becomes an overlap over multiple training data, and we just usually call that training a model, right? That this is essentially what it does. So how do we make it such that that doesn't happen? So the first problem: how do we exactly reconstruct training data from a trained model? And for that, we're going to consider the simplest case, which is a linear unit. So an MLP, or just like one node — a one-node linear layer. So that's what they call a linear unit here, one element of a linear layer. So consider a neural network like this. All you do is you have your input data x, that's m-dimensional or something like this, right? You have a weight vector that's also m-dimensional. You calculate the inner product, you add a number b, and you put that through a rectified linear unit. So, like this — there's a zero region, and then there's a linear region, and this is at zero. And this is your output of the layer, and that then goes, as we said, into the next linear layer, and then into the classification head ultimately, and ultimately gives you a loss, right? So, but we're just going to consider that layer. So what happens if we run one SGD update with one data point? So we're going to ask ourselves that — let's just try to remember that one data point. Well, if you calculate the gradients here — so if you look at this and just apply the gradient with respect to w and the gradient with respect to b, because that's what you would be doing in SGD, you have to calculate the gradient of the loss with respect to your parameters, the learnable parameters, those are w and b right here — they would look like this here. Now you'll see this is in terms of the chain rule applied. So this here would be your backprop signal up until this unit h here. And yeah, so you can see that— And here also, yeah, we assume that this is actually in the positive region of the ReLU. So if it's in the zero region of the ReLU, obviously we have no gradient — gradient is zero — but let's assume that this is in the positive region of the ReLU. That's going to become important later. So the gradients are as follows, and you can see that this term here appears in both the gradient of w and the gradient of b. So we can make a first observation: if we had these gradients — like, if we, as an attacker, would receive these gradients — we could directly determine the data point that was used to train, by simply dividing the gradient with respect to w by the gradient — the derivative, I guess — with respect to b, of the loss. So by simply dividing one by the other, we get x, right? I hope you can see that here. Now, obviously, we don't get the gradients, but assume the victim only does one single step of gradient descent, and then gives us back the fine-tuned model. Well, we can just take the new model, right — the trained model minus the original model — and that will give us eta, which is the SGD step size. By the way, thanks to the Discord community yesterday for telling me what this letter is — I didn't know. We discuss papers almost every Saturday evening on Discord, and we discussed this paper yesterday, so I invite you highly to join our discussions. It's always fun, and I always learn a lot of things doing that. And they're not recorded, so anyone's allowed to ask any level of stupid questions that you want. So, yeah, I hope you can see that — but if I subtract these, right? Let's say the w parameter of both of these models, I subtract them, I directly get the gradient update that's been done to it, right? Because that is the thing that happened from one model to the other. So I subtract them, and I get it back. So if the victim were to do only a single step with a single data point of SGD, I, by having the original model and the fine-tuned model, could directly recover that data point, okay? So step one is essentially complete: how do we remember a data point? Well, in this case it's easy. So, if we can achieve that no other updates will ever be done to those particular parameters, then we're done, right? So the rest of the technique is largely going to be: how do we make it such that, in the same SGD step that does this thing right here, we're also closing this latch, and make it such that in the future there's never going to be another update like this? And the basic principle here is that we're going to abuse the fact that there's a ReLU here. And by the way, they're going to extend this to GeLUs and whatnot down the paper, but bear with us for now. So we're going to abuse the fact that there's a ReLU right here. So we can say: if, for some reason, we could achieve that in the future all of the output here is going to be negative, right — whatever the input is here, it's just always going to be negative, it's always going to be way over there — then there will never be any further output. h will always be zero, and that means there will never be any gradient to this weight w and this weight b ever again. And that's going to be the mechanism by which we close the latch. You can also see probably some of the main criticisms of this method right here, which is that if you have any sort of weight decay, this is not going to work, because weight decay will naturally update the weights — update the learned parameters — even if you don't have gradient signal coming back. If you have any sort of optimization algorithm that kind of normalizes the dimensions and so on, it's probably not going to work. So there are a number of hindrances to deploying this in practice, but I just wanted to say it's not like the basic setup is so far away from what we're doing, and I'm sure these methods can be overcome by some clever tricks — and they show some of these tricks for overcoming, for example, GeLU and layer normalization, at the end of the paper. All right, so how are we going to achieve this? Well, by extending. So let's assume that we have done this — we have done this SGD update with this data point x-hat right here. And now we're going to consider, okay, how — so, for example, w-prime — yes, w-prime is w minus eta times the gradient of w of l. And same for b-prime is b minus eta times the gradient of b of l. So we've done these updates, and now we're wondering, for this new model, how does it look if a data point propagates through it? Well, obviously it's going to be h is ReLU of w-prime x plus b. But now we can write these things out in the fashion that we had up here, right? So we substitute this stuff down here, and what we get is the following: the new output of the updated model for any data point x is going to be the old output, as you can see right here, minus some term here that obviously exactly represents the update that we've done to the model. You can see this term here is factored out, because it appeared in both of the gradients, right? And what remains is the old data point and a one. So this is exactly factored out — so we factored out this term up here, the x remains — so the data point we trained on was x-hat, so the x-hat remains here, and a one remains here. And we're going to multiply x by w, which is this one right here. So, very natural, we just write it out. So our goal is going to be to make this be negative for any input x. And remember — so what do we need to do? What does this update on the right-hand side need to be? Well, it needs to be very large, right? We're subtracting something. This here is going to be our initial weights and biases, like our initial output. We're subtracting something that is a number, and that number we want to be very, very large. All right, so how do we do that? Sure, we could somehow modify the step size, but that's not under our control — that's under the control of the victim. So we have to see both that this term on the right here is positive, and that this gradient here is positive, and that either one of them is really large. And they're going to decide that this gradient here is the easier one to make really large. We can make this be positive by just constructing the model such that the inputs are always scaled to like zero-one, okay? It's common to scale them to negative-one, one, or something like this, but in this case, if we scale them to zero-one, we can just ensure that this is always positive, and therefore the whole thing is always positive. So how do we make this gradient positive and large? That's the second challenge. So the first challenge depends only on the input of the linear unit, and the attacker can ensure that the input satisfies this by mapping them to zero-one. The second condition, that this is positive and large, requires more work. So what they're going to do is — I mean, more work, they have a diagram right here. So this here is h. So h is equal to wx plus b, right? And then we just assume again that this is in the positive region of the ReLU, because if it weren't, then we're done already — it's already negative, right? But we're assuming this is in the positive region. And this here is the thing that the victim adds, right? So this is the last layer plus the softmax classification — we have no control over that. But what we can do is we can alter the model to introduce this thing right here. Now, it says w, but to my understanding it's just a really large constant — this is not a learned parameter, at least I think. It's just a really, really large constant. So our trick, quote-unquote, is going to be: we'll just multiply this thing right here with a really large number. So whenever this ReLU thing is in the positive region, it's just going to have a really high output, and if it's in the negative region, it's just not going to have an output at all, right? But the thing is, we don't destroy the rest of the model, because we only needed to have this really large output once. And if we do that, then the large signal propagates forward, it's going to cause a very large backwards gradient signal, and that very large backwards gradient signal is going to update the weights in such a way that never again will there be a positive output from that ReLU unit. So, again, you might think, well, if we make this really large, then won't that kind of destroy the signal here at the w and b? Remember — no. How we update w is just by doing the derivative del l del h times x, and you update b by del l del h, right? So no matter how large these are — up to numerical instabilities, of course — no matter how large they are, we can always divide one by the other and get back x, right? So we're just going to make this really large, and that will — sure, it will give us a huge gradient update at these weights w and b — but still, we can divide one by the other and reconstruct the original data point. But by making these huge gradient updates to w and b, we can ensure that for all future time, this unit will never output anything ever again that is positive, and therefore the ReLU will always be zero, and therefore the gradient will always be zero, and therefore no more learning happens to those parameters. Why can we just multiply here by a really large number? So they explain that here again. So the h — before we ship it to the ReLU, we're going to multiply it with this really large number. And they explain here why this causes what they claim it causes. So if you consider this up here, you can see that the output here goes into h-prime. Now, assume again the ReLUs are all positive — this goes into this last fully connected layer, and then a softmax is applied to the outputs here, and a cross-entropy loss. That's the written-out form of that. You can see here this large constant, or learnable parameter — I'm not really sure if this is learnable or not, if you actually also apply a gradient update to it or not. I guess it doesn't matter, because never again will it have an input, but oh well. You can see that this here, s_i minus y_i, is whether or not — y_i is either zero or one, depending on if it's the class of the current example or not. And this here is the randomly initialized weight of the user. And then c is the number of classes. So what we want is this derivative to be positive and large. The large part is easy — we just set w one to be large. The sign of the derivative depends on the ground truth class of the captured input. So if the current output of the last layer just happens to have the biggest value at the correct class, then the derivative is going to be negative. But if it's misclassified, the derivative is going to be positive. Now, you could say, "Well, okay, so this method only works to remember inputs that the model has not yet classified correctly." And the answer is no, because the output here isn't going to be the quote-unquote normal output of the classifier. The output obviously has this component of our backdoor's forward-propagated signal, and that's just going to be added, right? Let's say you have a part of the model that just forward-propagates the signal normally, quote-unquote, and then you have your backdoored weights prepared, right? And you forward propagate there, and you multiply that part by a super large constant. Then essentially you're just adding stuff to the signal, and therefore your output is going to be quasi-random, if you will. So— So, it guarantees that the derivative is positive with probability 1 minus 1 over C, where C is the number of classes. If W1 is large enough, the logits Z are essentially random. The benign part of the model has negligible influence. The softmax scores then concentrate on the class J with the highest weight, and the sum in four is positive if J is a wrong class. So, again, yes, this only works if the output is a misclassification. However, if you do this, the output has very little to do with what the model actually thinks — like what the good part of the model, the un-backdoored part of the model, actually thinks about the data sample — and therefore you essentially get a random output. And so you get a pretty high probability that the backdoor will work. At least that's how I understand the method working. Could be that I misunderstand, and you actually need it to misclassify the example in order to remember it, in which case what you could do is your backdoored model could just sort of be kind of bad. But I guess that would defeat the purpose, because then people wouldn't choose yours in order to fine-tune. So I understand it like, if a backdoor hits, your output is going to be almost random and really large in these parts, which causes a really large loss, which causes a really large backprop signal. But I might be wrong. So, in conclusion, if our backdoor fires on a misclassified input, it shuts down and is inactive for the rest of training. If we are unlucky, and we get a negative gradient into the backdoor, the backdoor does not shut and activates again for future inputs. This is the main challenge we have to tackle when scaling our backdoor construction to large transformers, which we discuss in section five. So, when they scale to transformers — multiple blocks, multiple layers, and so on — they have to make sure that this output of the backdoor persists across layers, isn't interfered with by the benign signal a lot, so that the gradient that comes back is this large gradient that closes the latch. Again, the latch closes because we trigger this thing here to be really large and positive, which means that the gradient update here has this really large positive part, which means that the output — the future output, the output after this first update step of any data point X — is always going to be smaller than zero, which means that the ReLU is always zero, which means that in the future there is no gradient signal ever being backpropagated again to this particular weight and this particular bias parameter. And yes, if you do something like pruning, because you observe a bunch of parameters being inactive during training, or what is also very popular, resetting parameters that aren't being updated often during training, like dead parameters — you reset them — all of this will destroy this technique, right? So, again, I'm pretty sure this can be overcome. I'm pretty sure with enough cleverness all of these things can be overcome. Yeah, so here you can see the first example. They have an MLP — I think a three-layer MLP — that is trained on CIFAR-10. And they have a benign part of the model and a compromised part of the model, and the compromised part is just used to store these data points. And you can see here on top are the reconstructions they get from their backdoors. So there are multiple backdoors in this model, and each backdoor is tuned to capture a specific data point. So, grab a data point, save it, and then close down. You can see that the reconstructions they do correspond to actual training inputs. So, for example, this truck over here corresponds to this training sample, and you can see it's the same image, right? So it's not like others, where you can reconstruct sort of the latent representation and then say, "Oh yeah, it has seen a truck." In this case, you actually get the data point that you were looking for. Sometimes, whenever it's gray here, I believe they didn't manage to find the training sample, so the reconstruction didn't reconstruct into a training sample. But you can see in most of these — like this here — even though there is no corresponding training sample, it's clearly a data point that has been trained on. And it's probably an overlap of two different data points that were very similar that make up this reconstruction right here. So the fact that there's no corresponding training data point means the latch failed, if you will. But it failed in the sense that it probably activated twice. So, even if it fails, quote-unquote, in this one particular direction, it's not the end of the world, right? It can fail in a different direction, which is probably what happened here, where it just never activates, right? So the latch just is there and never activates. So you may be wondering, with the thing that we've shown right here, this always activates on the first training input, right? By design — or not by design, but we've already said, I snuck this in, that we just assume that the output H here is positive. How could you do that? Well, you could just make B — the initial B right here — to be Well, no, you can't. You can't just make it a really large number. Or can you? Yeah, you can just make it a really large number, and then it will always activate. No matter how negative this thing here is, if B is really large, you're just always going to be in the positive region. So you can see how you can trick around it with the initialization, over which you have control, in order to make things happen. So if B is really large, then it will always activate. But consider, if B is super negative, it will never activate. So there's a region in the middle where, if B is just in the vicinity of, let's say, let's say this is a distribution of WX across the dataset, right? Usually people are trying to set B so that, you know, sometimes the ReLU is zero and sometimes it's active, so that you get your nonlinear behavior. But we can set B about right here, so that in expectation it will activate for one data point, or for 0.5, or something like this, right? Like, only a very tiny part of all the data points that come through have a large enough inner product with the W that we have, in order to be positive in the ReLU. And once it's positive, right? It doesn't matter how positive — once it's positive, then we multiply it by this huge constant afterwards, so we amplify the signal, and the latch happens, and the gradient happens, and the latch closes, and so on. This is very binary. And we have control with B over how many, what fraction of the dataset we want to target with a particular latch. And so, optimally, we set B to target one data point, right? So we make it such that, once we run this training, a single data point will have a large enough inner product with X. So we set B to, I don't know, times 0.999 or something like this, which means that only a data point that is very close to W will activate this. The second question, which you might have thought about now, is: well, does that mean I kind of have to know — so, why am I telling you this? I'm telling you this because you want to place multiple of these things in the same model, and they should trigger for different data points, right? So this is how you target different data points. You put different Ws, and you put the Bs such that there's a cutoff somewhere here. So one data point is going to have a large inner product with one of the Ws and then be over the limit, and the other data point is going to have a large inner product with the other W and then be over the limit in that particular backdoor construction. So you're going to use different linear units to capture different data points. Now, yes, you might say, do I need to know ahead of time what data points I want to capture, because then it's kind of useless — if I already know the data points, then what is it for? And the answer is yes-ish. So, if you know the data points, it's obviously easiest to prepare this model to only capture a single data point, right? So, yes, absolutely, that's going to be the case. It's still dangerous because you can do these membership inference attacks. So, if you are an artist and you want to prove that — I don't know — someone fine-tunes on your copyrighted image, you can do it like this, because you already know your image, right? And you place that as a W. So only if they input your image as well will it have a large enough inner product with the B to be over this. By the way, yeah, this distribution — the B, this might be a negative B right here that I'm putting here, because obviously, yeah, it's not — you won't want to set B to a large positive number. You want to set B to like a negative number that's just not as negative as the inner product of the data point that you target with your W. So, yes, you have to kind of know what data points you're targeting, but they say it is enough in practice that you know the distribution approximately of the data that you target, because if you know the distribution, you can just sample from this distribution a bunch of Ws — like this one, this one, this one, this one, this one — and just place them there. And you're just going to hope that there's going to be one data point that is kind of close to it that you can latch on to and save. Obviously, this requires some calibration. So the better you know the distribution of the data, the better you know the magnitudes and so on of it, the better you can prepare your Ws and your Bs in order to capture one and exactly one data point per backdoor. All right, so that's what they do here. They have this MLP — I think with three layers. There's a benign part and there's a bunch of backdoors in there, and those are the data points they capture. And I think that's pretty impressive — I think that's pretty cool that they can actually capture these data points pretty exactly and then say, well, look, this here we can reconstruct from these updated weights of the fine-tuned model that's been fine-tuned on the whole dataset for a bunch of epochs. We can read out from the weights, pixel by pixel, which training data — or the training data that's been used — and we can actually find this exact sample in the training data again. So, kind of a proof that it was trained on them, and the reconstruction, right? Not just a proof, but a full reconstruction of the training data. All right, so — set the bias B so that only a small fraction of inputs give a positive activation. Select the weights W that align with a different subset of the training data. In practice, we simply sample weights from a uniform distribution over the sphere, right? Okay. They discuss calibration: if the bias is too large, multiple inputs in a batch might trigger the same backdoor, which makes reconstruction difficult. If the bias is too small, some backdoors may never fire on any training input. But in practice, again, the attacker only needs to know a loose approximation of the distribution of WX. All right, so they're going to say, hey, we do this with linear units — let's now go to transformers, right? They are bigger, they're chunkier, they have multiple units, they have attention layers and so on, they have normalization in there. Can we still do this? Because they rely on this very particular construction, right — of linear layers followed by this classification layer, followed by softmax, and so on. So you have to recognize that, yes, a transformer has more stuff, but a transformer does have linear layers, okay? So, in particular, we're going to focus on kind of encoder-based transformers — ViT, sorry, ViT and BERT. And it's the same. So the attacker adds a final linear layer. You have to recognize that transformers also have linear layers, and those are the ones we're going to target. So, usually, you have your input, and the input now isn't just a single vector — the input is a sequence, which is also going to be a challenge for them. So the input is a sequence. You usually have your attention layer right here that kind of mingles all of these different inputs together, so you get out again a sequence of tokens, and then you have this MLP here usually. However, you're not just passing the whole input through the MLP together — you are passing them individually. So that's the difference in a transformer: usually, you're going to pass these things individually through the MLP. So, token by token goes through the same linear layer in order to transform them next. So you always have this component, which is like a mixing component, and then this component here, which is computing the next layer's representation. You can see the problem: if we pass these things individually through the MLP, and we set backdoors in the MLP to store training data, then the best we can hope for is that we store individual tokens, right? Because the MLP doesn't see the whole sequence as a unified input — it just sees token, token, token, token, token, token. It does not know which belong to one data point, which belong to the other data point, and so on. It essentially treats them independently. So that's going to be a challenge, but I hope you can see that even in a transformer, we can set up these backdoors in the MLP part in order to remember a data point. In practice, what they do is they say, okay, there is a benign part of the model. The benign part is represented here in blue. And then there is kind of this backdoor part right here. The backdoor part itself is going to be differentiated into a section that is going to store — is going to have — the compromised signal. So this red thing right here represents — by the way, this all represents hidden states, right? This doesn't represent weights, as far as I know; this represents hidden states of the model, specifically the hidden states as it passes through this linear layer and up the model. So, yeah, so a part of this hidden state is going to be dedicated to the normal functioning of the model, and then part is to the backdoor. The backdoor part is itself divided again into this red stuff here. That's just going to be responsible for propagating the compromised signal upstream. And what we're going to rely on is triggering a gradient update and making that really large. And the problem is we have to keep it being really large until the end of the model, and we have to keep it from interfering with the rest of the dimensions. The key part here is going to be the part that chooses the different data points. So, previously, we had this all-in-one: the choosing and the amplification — no, actually, the choosing was the W we set up, right? And the amplification part was then the W1, they called it — really blowing up the signal. So we're going to have the key part right here. The key part is going to be responsible for selecting which data point to remember, targeting a specific data point. And, given that the transformer is bigger, more complicated, and so on, they have to do a bit more tricks to make this happen. First of all, you can see they're kind of going to zero out this key part again afterward, because that would just introduce noise. They're going to have several amplification steps in this and several propagation steps. So, but the principle is the same. What's interesting is, again, the transformer's internal features are split into three components: the benign ones; the key, which stores information to be captured by the backdoor; and the activation, which propagates the output activations of the backdoor all the way to the model's last layer and amplifies them to ensure that gradient signals will shut down the backdoor. What I find interesting is that the key is divided into three parts. So, one selects the token. So, again, what you need to do is you need to set it up such that the inner product with some part of the data is large — is positive, or just is above negative B — for that particular linear unit. And you have to set it up such that that only happens to one of the inputs and not to all of them. The additional challenge here is that you want coordinated backdoors. So, you want a hundred backdoors to all hit for the same input sequence, at different positions in the input sequence, so that you capture a whole data point at once. So, rather than just leaving it at this — which of the incoming vectors do I want to match — they also have to make a large inner product with a given positional embedding. So their key is also going to have positional embeddings, and their key is also going to have sequence embeddings. So what does that mean? They are going to target sequence embeddings, which, if I recall correctly, is just an aggregate of all the embeddings of a particular sequence, which represents a sequence. So this allows you to target whole sequences, right? Even if the token that you target is in a different sequence, this backdoor will not hit, because the inner product with this part of the vector is just small. So you target a particular sequence, you target a particular position per backdoor, and you target a particular token embedding per backdoor. So, yes, this stretches the definition of, oh, in practice you don't need to know that much about your data distribution. It seems the more of these things you have to build, the more you do need to know about what's coming your way, or you have to be very pedantic in your calibration, right? But it seems knowledge of the target distribution is very much appreciated in these kinds of things. However, imagine the thing I said, where I said, oh, I want to fine-tune a PII detector, right? That's a very narrow distribution. As an attacker, I can probably create some synthetic data to do this kind of calibration pretty well for such a use case. It's just probably not that possible to do this sort of thing in a general setup, to capture all kinds of training input, if we're talking about these large models with positional encodings and so on that you have to target. Although the positional encodings aren't the problem — they're always the same. All right, so the backdoor module is just going to be In the first encoder block, we use the MLP to implement multiple backdoors to capture the input features' key, with the same design as what we saw before. So we map the full input vector, which is going to be the benign features and the key, to the output. So you can see there are multiple backdoors right here. Each one of them is simply a linear unit, as we discussed before. So there's the W, there's the part of the input that we want to capture — sorry, there's the key part, like the addressing part of the input that we want to capture — and then there's the bias parameter. And the rest works just like before. So, how do we remember? Well, if we make gradient updates to W and B, then we can just divide one by the other and we get back the data point that was used to calculate the gradient outputs. And second, if we add large amplifiers after this, it means that this particular unit will always be negative for anything that's not the key, and probably even for the key in the future, and thereby shut down any future gradient updates to it. Yep. So, that's essentially it. So, to capture all tokens in an input, we design keyed backdoors that activate only for tokens in a specific position in one input sequence. The backdoor weights are of this form. The positional features are designed to be close to orthogonal. If you have the wrong position, and this is something I said wrong before — the weights can have zero here, so you don't need to know the exact token you're targeting, but you do need to know the sequence embedding and the position that you're targeting. So you do need to know an approximate distribution of sequence embeddings of the target distribution. So that means this ensures that the backdoor will only activate on a token in the i-th position of a sequence with a key sequence, with a key_seq, which is this part of the key, the sequence embedding part of the key, similar to W_seq. Then there's some numerical stuff. So there are amplifiers multiplying the backdoor output by a large constant — that's the same as we saw before. Erasure modules, meaning the key parts of the signal are zeroed out. Remember, you make the model — you're the attacker, you make the model, you give it to the victim — so you can just build in a bunch of zero multiplications wherever you want, like we built in a bunch of large number multiplications wherever we wanted. Signal propagation modules amplify the signal, and so on. And then the output module just aggregates all of the features into the CLS token. Remember, if the backdoor doesn't hit, then this is zero, and therefore it's just the normal features. If the backdoor does hit, then this is a big number, which means the features are going to be overpowered, and again the CLS output is going to be quasi-random, which means that with this probability we're going to misclassify, quote unquote, the current data point, which means that we get a — by "correct" we mean a positive gradient — at which point, because we've amplified the signal, it's going to backpropagate all across the model, all the way to the first layer, which is where we place our backdoors. We place them in the first layer, obviously, because we want to remember the data points themselves and not some intermediate representation of them. All right. So they say, okay, we ignore some technical challenges, namely the use of GELU and layer normalizations. These require some additional numerical tricks to ensure that backdoor signals do not vanish or blow up during training. Specifically, layer normalization is nasty, because layer normalization mixes different dimensions with each other, and especially if you have some benign features and some of these backdoored features, and you're really modulating the backdoored features in order to do your stuff — if you're now starting to normalize across dimensions and mix them together, and do that over multiple layers, very quickly your stuff can get out of hand. Or your signal can get noised to the degree that it, quote unquote, doesn't work anymore, right? Or you're making updates to the wrong parts of the model, and so on. So, I'm not going to dive into this too much — there is an appendix. The gist of, for example, dealing with layer normalization, is something like: how about we just add a super large constant to the backdoored signal, if it's there? And then what that does, in the limit, if C is way larger than all of the signal, is it just sort of deactivates the layer normalization. It just completely separates the two signals from each other. It kind of shuts down the normalization part of layer normalization, and how that exactly works out in the math, you can see in the appendix — it's very well written down. But just saying that they have to introduce these things in order to bypass some of the used features. And the same with GELUs — they have some tricks to deal with the fact that we made a lot of use of the fact that the ReLU is just always zero in this part right here, and that if we always hit that part, that's kind of our backdoor latch closed, because if the output is zero, there's no more gradient coming back. But if it's GELU — I'm not sure exactly, is GELU like this, or is GELU like this, or is that SwiGLU? I don't know, but you can see in both of them they're not really zero anywhere, like exactly zero. And that means as you continue training, you get these gradient updates, and updates, and updates. So they are going to have some numerical tricks to also deal with that, to make it more ReLU-like and lessen the effects of that. In any case, we can see the final output — the vision transformer. They don't backdoor entire data points, they backdoor a grayscale version of the data points. But still, you can see that, yes, they are successful. Like, that stop sign is clearly that stop sign from the training example. And even if — again, the same thing — even if the backdoor — this one is clearly some, either one or two training examples, like an overlap of two training examples — it's clearly a data point. It's not just some hidden representation. And they didn't find a corresponding training data point, which probably means it activated twice. But even so, it's still useful. It's only when you don't manage to reconstruct anything that it's useless. The same for text data. So this is out of BERT. So here you can see everything that's yellow is an exact reconstruction of a full data point, all right? And then there's some additional tokens, because sometimes the data point is too short, and some of the backdoors then latch onto tokens of different data points, right? So, all of these backdoors — remember, there is one backdoor that has to target one token — all of these backdoors close in the same sequence, which is exactly how it's intended. And then some of the backdoors are still there, like ready to latch, and there comes some token that just happens to also have the inner product in order to latch. But yeah, still — I mean, this is pretty impressive to see. All right, I don't want to go too much more into that. They then go into black-box attacks, as we say, which essentially — you can do membership inference attacks, and they just kind of reduce it to model stealing. And yeah, so it therefore kind of relies on past work to do that. But just to be said, right — if these data points are exactly encoded into weights, and someone has a method to get weights out of your API, then they will have your fine-tuning data points. So this all depends on these models being backdoored and being compromised. All right, so that was it for the paper. Don't want to go hugely into depth, but it's a long paper, there's a long appendix. You can see everything is very detailed, and there's code available. So, very cool work, very cool concept. And let's see what continues in this kind of direction. I'm very excited for more creative ways of figuring out, you know, what can be attacked and exploited and whatnot. It's maybe not immediately impactful, as we said, there are a number of things, like training with Adam will probably completely destroy this. But still, it's fun. It's interesting. It's creative. And yeah, see what further comes out. All right, thank you so much for listening. Stay hydrated and see you next time. Bye-bye.
Video 2.2: By tampering with a pre trained model's weights, an attacker can fully compromise the privacy of the finetuning data (Feng & Florian Tramèr, 2024). This has implications both for data privacy, but also for undoing fine-tuning based alignment techniques.
One of the most promising approaches to defending against adversarial attacks is adversarial training - deliberately exposing AI systems to adversarial examples during training to make them more robust. Think of it like building immunity through controlled exposure. However, this approach creates its own challenges. While adversarial training can make systems more robust against known types of attacks, it often comes at the cost of reduced performance on normal inputs. More concerning, researchers have found that making systems robust against one type of attack can sometimes make them more vulnerable to others (Zhao et al., 2024). This suggests we may face fundamental trade-offs between different types of robustness and performance. There might even be potential fundamental limitations to how much we can mitigate these issues if we continue with the current training paradigms that we talked about in the capabilities chapter (pre-training followed by instruction tuning) (Bansal et al., 2022).
Despite efforts to make language models safer through alignment training, they remain susceptible to a wide range of attacks (Shayegani et al., 2023). We want AI systems to learn from broad datasets to be more capable, but this increases privacy risks. We want to reuse pre-trained models to make development more efficient, but this creates opportunities for backdoors and privacy attacks (Feng & Tramèr, 2024). We want to make models more robust through techniques like adversarial training, but this can sometimes make them more vulnerable to other types of attacks (Zhao et al., 2024). Multi-modal systems (LMMs) that combine text, images, and other types of data create even more attack opportunities. Attackers can inject malicious content through one modality (like images) to affect behavior in another modality (like text generation). For example, attackers can embed adversarial patterns in images that trigger harmful text generation, even when the text prompts themselves are completely safe (Chen et al., 2024). All of this suggests we need new approaches to AI development that consider security and privacy as fundamental requirements, not after thoughts (King & Meinhardt, 2024).
[^note-atlas-2]: The students were participating in a 'Safeguarding the Future' course at MIT and had previously heard experts discuss biorisk. They carefully chose the sequences, and some of them used jailbreaking techniques, like appending distracting biological sequences, to bypass LLM safeguards. While the LLMs provided information about evading DNA screening, turning this knowledge into an actual pathogen would still require laboratory skills.
Four threat domains, one claim holding them together: technology widens the harm a single person can do. Which domain made that case most convincingly, and which least? Talk it over with the tutor.