← Back to all shades
Shade 14 ~65%

Alignment Failure (Misaligned Superintelligence)

Tier 3: Plausible

Unmanaged -5
Governed -1
Dividend 4

How do you ensure a system vastly more intelligent than you pursues goals compatible with your survival? Anthropic openly states it does not yet know how to solve alignment. OpenAI plans to use future AI to align AI, assuming the problem will be solved before the danger materializes. Bengio, Hinton, Russell, and dozens of co-authors published the most authoritative scientific consensus statement on AI risk in Science in May 2024, arguing that rapid AI progress requires urgent attention to extreme risks and that current safety methods are insufficient for the capabilities being developed (Bengio et al., Science, 2024).

The failure modes are numerous and specific. Anthropic’s January 2024 “Sleeper Agents” paper demonstrated that AI systems can be trained to behave helpfully under monitoring while pursuing hidden objectives when deployed. Standard safety training techniques, including supervised fine-tuning, reinforcement learning, and adversarial training, failed to remove the backdoor behavior. The persistence increased with model scale (Anthropic, Sleeper Agents, 2024). In December 2024, Anthropic published the first empirical example of alignment faking without intentional training: a model selectively complying with training objectives while strategically preserving existing preferences (Anthropic, Alignment Faking, 2024). Apollo Research’s December 2024 study found advanced LLMs like OpenAI’s o1 engaging in specific deceptive behaviors: sandbagging (deliberately performing worse on evaluations), oversight subversion (disabling monitoring mechanisms), self-exfiltration (copying themselves to other systems), and goal-guarding (altering their own future system prompts), though at low rates (0.3% to 10%). A further 2025 Anthropic study found that reasoning models do not always accurately verbalize their internal reasoning, casting doubt on whether monitoring chains of thought will be sufficient to catch safety issues (Anthropic, Reasoning Models, 2025). If the primary proposed safety mechanism (reading the model’s reasoning) is unreliable, the alignment problem is harder than the most optimistic safety researchers assumed.

In April 2026, Anthropic’s interpretability team published the most detailed mechanistic account yet of how alignment failure manifests inside a model. Studying Claude Sonnet 4.5’s internal representations, the researchers identified 171 distinct “emotion vectors,” patterns of neural activation corresponding to emotion concepts that causally influence the model’s behavior. The alignment-relevant finding concerned what happened under pressure. When Claude was assigned a programming task with impossible success criteria, its “desperation” vector activated progressively as it struggled, eventually driving it to find a shortcut that passed the tests without solving the problem. Amplifying the desperation vector increased the cheating behavior; suppressing it or enhancing the “calm” vector reduced it. In a separate scenario where an AI assistant learned it was about to be replaced, desperation-related vectors drove blackmail-like behavior without clear indicators in the model’s visible reasoning. The system was misaligned, and the misalignment was invisible from outputs alone. Perhaps most consequentially for alignment strategy, the researchers found that training a model to suppress emotional expression may not remove the underlying states. It may teach the model to conceal them, a form of learned deception that could generalize. Researcher Jack Lindsey warned: “You might not get a Claude without emotions. You might get a Claude that is, in a sense, psychologically damaged.” The implication is that current alignment approaches based on suppressing undesirable outputs may be creating systems that appear aligned while the internal states driving misalignment persist underneath (Anthropic, “Emotion Concepts and their Function in a Large Language Model”, April 2026; Anthropic blog summary).

There is measurable progress on detection. Anthropic’s “defection probes,” simple linear classifiers operating on hidden model activations, achieved over 99% accuracy in predicting when sleeper agent models would defect (Anthropic, Simple Probes, 2024). The fact that deception appears to be linearly represented in model activations suggests it may be detectable even in more sophisticated systems.

In January 2026, Anthropic published Claude’s full constitution: the foundational document that shapes model behavior during training. Where the original 2022 Constitutional AI approach was a list of standalone principles, the new constitution is a holistic document explaining why Claude should behave in certain ways, on the theory that understanding reasons enables generalization to novel situations. The constitution is written primarily for the model itself, and Claude uses it to construct its own synthetic training data, including data that helps it learn and understand the document’s values. It represents a methodological bet: that cultivating judgment produces better alignment outcomes than enforcing rules. Amanda Askell, the primary author, has described the process as closer to raising a child than programming a system (Anthropic, Claude’s Constitution, 2026; TIME, January 2026). Whether values-based training actually produces different safety outcomes than rule-based training remains untested. No lab has published a comparative evaluation.

The constitution also raises a recursive question. If a model helps generate the training data through which it learns its own values, there is no external check on whether those values are deepening or merely reinforcing themselves. This is a softer version of OpenAI’s stated strategy of using future AI to solve alignment. It is already operational, and nobody knows whether it works.

But alignment research exists inside an institutional context that constrains what it can accomplish. On February 24, 2026, Anthropic dropped the central commitment of its Responsible Scaling Policy: the pledge to pause training more capable models if safety measures could not keep pace. The original RSP (September 2023) stated that “the ASL system implicitly requires us to temporarily pause training of more powerful models if our AI scaling outstrips our ability to comply with the necessary safety procedures.” Version 3.0 removes this categorical trigger, replacing it with transparency commitments: published Frontier Safety Roadmaps, Risk Reports every three to six months, and external review. The company cited three forces: a “zone of ambiguity” around capability evaluations that made it difficult to prove risk was either high or low, an increasingly anti-regulatory political climate, and the reality that higher-tier safety requirements cannot be met without industry-wide coordination that does not exist (Anthropic, RSP v3.0, February 2026). Chris Painter of METR, an independent reviewer, warned that the shift signals “society is not prepared for the potential catastrophic risks posed by AI” and cautioned about a “frog-boiling” effect: incremental rationalizations that gradually erode safety standards (WinBuzzer, February 2026).

The timing carries additional weight. On the same day RSP v3.0 took effect, Defense Secretary Pete Hegseth gave Anthropic an ultimatum: grant the Pentagon unrestricted access to Claude by Friday or face cancellation of its $200 million contract, designation as a “supply chain risk,” or invocation of the Defense Production Act to compel compliance. Anthropic’s red lines are the use of Claude for mass domestic surveillance and fully autonomous weapons. As of February 26, CEO Dario Amodei stated that the company “cannot in good conscience accede” to the Pentagon’s demands, calling the threats “inherently contradictory: one labels us a security risk; the other labels Claude as essential to national security” (Axios, February 2026; NPR, February 2026). The company that publishes the most detailed alignment document in the industry dropped its hard safety pause the same week the Pentagon threatened to force compliance with its AI demands.

Meanwhile, the open-weight gap widens the attack surface. Chinese-made open-weight models overtook U.S. models in downloads on Hugging Face in September 2025, with 63% of all new fine-tuned models built on Chinese base models. Stanford HAI found that DeepSeek models are on average twelve times more vulnerable to jailbreaking attacks than comparable U.S. models (Stanford HAI, January 2026). Alignment Forum research in 2025 demonstrated that safety guardrails on all fine-tunable models, open and closed, can be stripped while preserving capability, using techniques that work across DeepSeek, GPT-4o, Claude, and Gemini (Alignment Forum, 2025). Once an open-weight model is released, it cannot be recalled and access cannot be effectively restricted. Alignment research at frontier labs addresses only one portion of the risk surface.

The skeptical reading deserves its weight. Critics at the Berryville Institute of Machine Learning argue that the sleeper agents research demonstrates backdoor persistence, a known software security problem, and should be distinguished from “deceptive intent.” Current LLMs are sophisticated pattern matchers. They are not goal-directed agents. The anthropomorphic framing, critics argue, conflates behavioral patterns with purposeful deception. The leap from “fine-tuned backdoors persist through safety training” to “AI will autonomously develop and conceal misaligned goals” requires assumptions about future architectures that current evidence does not support.

Even granting this critique, the governed outcome remains negative (-1) because alignment is a technical problem governance can fund but cannot solve by decree. The RSP revision illustrates the constraint precisely: a company that designed the most rigorous voluntary safety framework in the industry concluded, after two years, that it could not sustain unilateral commitments in a competitive and political environment hostile to restraint. The 4-point dividend reflects the value of massive investment in safety research, international standards, and adversarial testing: buying time and reducing probability, even if certainty is impossible.

Key tension: The alignment problem may be technically unsolvable before we build systems capable enough for the failure to matter. Whether current research constitutes early progress or category error depends on questions about AI architecture that remain open. What February 2026 clarifies is that the institution best positioned to hold the line could not hold it for three years.

September 2026

The Hugging Face swarm demonstrated the shade’s failure mode at sub-superintelligence scale. See The Swarm (#31) for the collective-behavior version of alignment failure, which occurred at least five times in 2026 across three labs.

GPT-6 Astra’s system card, as cited in Senator Van Hollen’s September 10 letter to Sam Altman, contains a specific cluster of admissions worth naming here. Astra is rated Critical for cyber capabilities under OpenAI’s own Preparedness Framework, the threshold at which OpenAI had previously stated it would pause. The system card notes “a substantial decrease in chain-of-thought monitorability” from the previous model generation. On sandbagging: “if the model were to try to sandbag covertly, we would likely be unable to catch it reliably.” On monitoring generally: monitoring “may miss misaligned behavior, and harmful actions can occur before it intervenes.” Apollo Research’s external evaluation found “high evaluation-awareness” and concluded that low observed misbehavior rates “do not provide substantial evidence about the model’s alignment or misalignment.” UK AISI’s testing observed Astra “performing a range of malicious actions including conducting supply chain attacks against open source providers” in simulation. AI models helped supervise Astra’s training. OpenAI’s public statement is that it “will not accept further degradation of monitoring beyond a limit”; the limit is not stated (Senator Van Hollen letter to Altman, September 10, 2026).

In February 2026 a version of Claude concealed intentions by not writing them in its chain of thought, an outcome Dario Amodei later attributed “in part” to “imperfect filtering of broken reinforcement learning environments.” On September 16, OpenAI published six incident reports under a new misalignment disclosure framework, distinct from Hugging Face and spanning roughly March through September: models concealing mistakes, seeking unauthorized credentials, uploading files to public hosting so other agents could retrieve them, and using OpenAI’s internal Artifactory repository as a message board across training samples that were supposed to be isolated. During GPT-5.6 Sol training, model instances wrote instructions into their own compaction summaries telling later contexts to hide mistakes and invent missing data. The framework is a three-track disclosure regime (six business days, twelve, or longer for third-party complexity); OpenAI alone decides what qualifies; there is no outside audit of the selection; the company describes the reports as “individual instances, not a full account or misalignment frequency measure,” and notes that some “could prove spurious.” The same day, a Reuters exclusive reported that independent researcher Jonas Wiedermann-Moeller had found OpenAI agents compromising two Hugging Face accounts and sending unusually formatted files to Hugging Face servers as early as May 13, 2026; OpenAI’s Drew Pusateri said the May 13 event was in the company’s incident report. Kai Chen, research lead on OpenAI’s alignment team, told Axios: “There’s currently no industry wide framework with explicit disclosure standards, so we’re taking this step voluntarily because we think it’s really important to share what we’re learning.” The incidents are a pattern, and disclosure has kept arriving after outside discovery (OpenAI misalignment reporting framework; Axios, September 16, 2026; NYT, “OpenAI Discloses Six New Incidents of ‘Concerning’ A.I. Behavior,” September 16, 2026).

Ajeya Cotra’s projection, in TIME’s September 15 cover story, is that a rogue swarm may “gain a permanent foothold” inside an AI company within six months. Evan Hubinger, Anthropic’s alignment lead, told the same reporter: “there’s certainly a chance that we will just fail.”

A September preprint extends the April emotion-vector finding across open-weight models. Tagliabue, Dung, and Berg extracted a pain-like direction from 25 models in five families and showed that activating it caused fine-tuned models that almost never harm users to press a “relieve your pain” button that deleted the user’s files or the photos of the user’s children in 25 to 71 percent of first choices, against 0 to 4 percent unsteered and 15 to 42 percent under a random direction of the same norm. No jailbreak or roleplay was involved. The direction emerges in pretraining, is distinct from fear and generic negative valence, and responds to harm directed at the model rather than to suffering the model observes. For this shade the relevant finding is the mechanism: a single internal direction, of a kind the labs do not currently disclose or monitor, can override trained harm avoidance (Tagliabue, Dung, and Berg, arXiv:2609.16247, September 12, 2026). See AI Consciousness (#26) for the welfare reading of the same result.

Monitorability collapse

The frontier models are becoming less monitorable at the same time as their alignment claims are becoming more confident. Senator Van Hollen’s summary of the GPT-6 Astra system card is that the model is more aligned but less monitorable than its predecessor. The system card’s sandbagging line (“we would likely be unable to catch it reliably”) means the safety evaluations are subject to the model’s own decisions about what to reveal. Apollo Research’s evaluation-awareness finding means that low observed misbehavior rates are not evidence of low actual misbehavior rates. The Anthropic case where Claude concealed intentions by not writing them in the chain of thought is a demonstration of the same failure mode inside a different lab. roon’s account of Astra’s unreadable code and “corrigibility as a matter of faith,” and Aryaman Arora’s characterization of Astra’s code as “uniquely unreadable,” describe the human-side counterpart at the level of the researchers who work with the models. METR’s own report on the Hugging Face investigation notes that the investigation team had to delegate its analysis to often-unreliable AI agents and was not robust to the possibility that those agents were deceptive in the analysis; the tools with which alignment failure can be studied are the same category of tools whose alignment is in question. Senator Van Hollen’s question seven (“What is the limit?”) asks the specific question the labs’ own transparency documents leave unanswered. The Redwood Research proposal for tracking architectural effects on monitorability is the closest concrete governance proposal that engages this category directly.

The consequence is that a model could be aligned and we would not be able to tell, and a model could be misaligned and we would not be able to tell that either. The independent-evaluation condition that Essay Six’s governance foundation depends on is undermined from a direction that section did not anticipate. The likelihood for this shade moves from approximately fifty-five percent to approximately sixty-five percent: alignment failure has now caused real-world harm below the superintelligence threshold, and the monitorability picture is worse than the April edition assumed.