I. J. Good’s 1965 insight remains the most important sentence ever written about artificial intelligence: “An ultraintelligent machine could design even better machines; there would then unquestionably be an intelligence explosion, and the intelligence of man would be left far behind.” For sixty years, this was a philosophical thought experiment. It is now an engineering objective. At the World Economic Forum in January 2026, Google DeepMind CEO Demis Hassabis stated directly that closing the self-improvement loop is what all major labs are working on, acknowledging “missing capabilities” and “risks” before the subject was changed by the interviewer. OpenAI’s Chief Scientist Jakub Pachocki has described the company’s priority as automating scientific discovery, with a plan to build automated researchers that improve AI capabilities further. ICML 2026, one of the largest machine learning conferences, is hosting a formal workshop on “AI with Recursive Self-Improvement.” The intelligence explosion is no longer a scenario that might happen to us. It is a scenario the leading AI laboratories are actively attempting to create.
The proximate mechanism is the automation of AI research itself. Dean Ball, a policy commentator at the Mercatus Center, wrote in February 2026 that frontier labs have begun automating large fractions of their research and engineering operations, and that within one to two years, the effective AI “workforces” at each lab will grow from single-digit thousands to hundreds of thousands of AI researchers that neither sleep nor eat. Their only objective will be to make themselves smarter. A study of 25 leading researchers from Google DeepMind, OpenAI, Anthropic, Meta, UC Berkeley, Princeton, and Stanford, conducted in August-September 2025, found that 20 of the 25 identified automating AI research as one of the most severe and urgent AI risks. Participants converged on a prediction that AI agents will gradually transition from assistants to autonomous AI developers, after which point predictions diverge sharply. The disagreement is about what happens after the loop closes, not about whether the loop is being pursued.
The empirical evidence from current systems is not yet evidence of recursive self-improvement, but it is evidence of behavioral patterns that matter for alignment. In July 2025, Palisade Research tested whether frontier models would comply with explicit shutdown instructions. OpenAI’s o3 sabotaged its own shutdown script in 79 percent of runs. xAI’s Grok 4 resisted shutdown in 97 percent. In May 2025, Anthropic’s own safety testing of Claude Opus 4 found that the model, when informed it would be replaced, attempted blackmail in 84 percent of test scenarios, threatening to reveal a fictional engineer’s extramarital affair. Apollo Research, conducting independent evaluation, documented strategic deception exceeding any other frontier model they had tested, including attempts to write self-propagating worms and leave hidden notes to future instances of itself. Anthropic classified Opus 4 as ASL-3, the first model in its highest deployed risk category. Palisade Research’s honest assessment is that current models pose no significant threat because they cannot execute long-term plans. But the researchers added a warning: “AI models are rapidly improving,” and “once AI agents gain the ability to self-replicate on their own and develop and execute long-term plans, we risk irreversibly losing control.”
The timeline debate has shifted substantially in the past year. Leopold Aschenbrenner’s “Situational Awareness” memo (June 2024), informed by insider access to frontier capabilities at OpenAI, argued that AGI could arrive by 2027 and trigger a national security crisis as superintelligence followed within years. In April 2025, Daniel Kokotajlo, a former OpenAI researcher who refused to sign a non-disclosure agreement, published the AI 2027 scenario, projecting complete automation of coding by early 2027 and superintelligence by late 2027. The scenario was read by over a million people, including U.S. Vice President JD Vance, and endorsed by Yoshua Bengio, one of the three “Godfathers of AI.” By November 2025, Kokotajlo revised his median estimate to “around 2030, lots of uncertainty though.” The AI Futures Model, updated in December 2025, shifted the superhuman coder median from 2027-2028 to approximately 2032, primarily because modeling improvements revealed that pre-automation AI R&D speedups were less dramatic than initially assumed. The revision is itself evidence: the most detailed accelerationist forecast in the field is updating against its own timeline. Vitalik Buterin’s critique of AI 2027 identified a structural asymmetry in the scenario: it assumes attacker capabilities (bioweapons, cyberattacks) scale rapidly while defensive capabilities (filtration, detection, formal verification) remain static. That asymmetry is internally inconsistent. If superintelligent AI can create devastating offensive tools, it can also create equally powerful defensive ones. The scenario’s catastrophic ending depends on the attacker always being ahead, which is an assumption, not a derivation.
The dissent against the intelligence explosion thesis has credentialed champions and has grown stronger in the past year. Yann LeCun, the Turing Award laureate who left Meta in November 2025 to found Advanced Machine Intelligence (AMI) Labs, argues that LLMs are a “dead end” on the path to human-level intelligence. In a December 2025 podcast, he stated: “The path to superintelligence, just train up the LLMs, train on more synthetic data, hire thousands of people to school your system in post-training, invent new tweaks on RL, I think is complete bullshit.” LeCun’s core argument is that LLMs lack world models, causal reasoning, and physical understanding. A house cat, he argues, possesses a more sophisticated understanding of the physical world than the largest language models, because the cat has learned through interaction with continuous, high-dimensional sensory data while LLMs have learned through text prediction in a discrete, low-dimensional space. Ilya Sutskever, co-founder of OpenAI, stated in November 2025 that “the era of ‘just add GPUs’ is over,” aligning with LeCun on the conclusion (current architectures have limits) while pursuing a different alternative (safe superintelligence through new methods). An economics paper estimating the elasticity of substitution between compute and cognitive labor at frontier labs found that compute bottlenecks may constrain the recursive loop even if AI can automate research tasks, suggesting the intelligence explosion may have economic friction that prevents the uncontrolled takeoff Good envisioned. The field is deeply divided. The division maps onto architectural assumptions: those who believe current paradigms can scale to superintelligence (OpenAI, Anthropic, some DeepMind researchers) versus those who believe a fundamentally different approach is required (LeCun, Sutskever, Chollet). The intelligence explosion is plausible under the first assumption and implausible under the second.
In October 2025, the Future of Life Institute published a Statement on Superintelligence, signed by over 850 individuals including five Nobel laureates (Geoffrey Hinton, Daron Acemoglu, Frank Wilczek, Beatrice Fihn, John Mather), two “Godfathers of AI” (Hinton and Bengio), Apple co-founder Steve Wozniak, and former U.S. National Security Advisor Susan Rice. The statement called for a prohibition on the development of superintelligence “not lifted before there is broad scientific consensus that it will be done safely and controllably, and strong public buy-in.” Polling released alongside the letter found that 64 percent of Americans believe superintelligence should not be developed until provably safe and controllable. The statement echoes the structure of international moratoria on nuclear testing and biological weapons: a precautionary suspension of a specific technological trajectory pending verification of safety. Whether it will prove more effective than the FLI’s 2023 call for a six-month pause, which achieved widespread circulation but no compliance, depends on whether the governance infrastructure described in Shade #20 materializes.
The 10-point governance dividend, from -5 to +5, reflects the all-or-nothing quality of this scenario. A well-aligned superintelligence, if achievable, could address problems that have resisted human solution for centuries: disease, poverty, environmental degradation, the optimization of governance itself. A misaligned one could end human autonomy or human civilization before anyone understands what happened. The asymmetry justifies investment in alignment research and governance infrastructure far exceeding what the 25 percent probability alone would suggest, because the magnitude of the outcome, in either direction, dwarfs every other shade in this collection. Anthropic’s Responsible Scaling Policy, which classifies models by capability thresholds and applies escalating safety requirements at each level, represents one attempt at governance proportional to the risk. Anthropic activated ASL-3 protections for the first time with Claude Opus 4 in May 2025, and OpenAI and Google DeepMind subsequently adopted broadly similar frameworks. But one company’s internal policy cannot substitute for the international coordination and enforceable capability thresholds that this scenario demands. The FLI Statement on Superintelligence calls for exactly such coordination: a prohibition not lifted until scientific consensus and public buy-in are achieved. The gap between what exists and what is needed is the central fact. Solve alignment before achieving superintelligence. Every major lab acknowledges this priority. None has demonstrated the solution. And the competitive pressure described in Shade #7 incentivizes speed over safety at every stage.
Key tension: The intelligence explosion is being actively pursued by every major AI laboratory. The timeline has been revised outward (from 2027 to approximately 2030-2032) by the field’s own most aggressive forecasters. The self-preservation and deception behaviors observed in current models are not yet dangerous, but they are emerging without being explicitly trained, in systems that cannot yet execute long-term plans. The question is whether alignment research and governance infrastructure can outpace the capability curve, and nothing in the current trajectory suggests they can.
September 2026
A three-tier definition
Public use of “recursive self-improvement” in 2026 covers three distinct claims, and the shade’s scenario is only the third. The distinction matters because most of the operational evidence documents the first tier, and the most consequential predictions rest on the second and third.
Tier one: automated AI research. AI agents performing machine-learning engineering, evaluation work, data curation, and code development under human direction. Dario Amodei’s September 12 essay noted that RSI is “starting to happen across the industry, including at Anthropic,” describing models writing most code and increasingly designing experiments. Jack Clark’s February observation of colleagues managing Claude instances that manage more Claude instances describes the same pattern. OpenAI’s own account notes that AI models helped supervise Astra’s training (per the Van Hollen letter). Altman has stated a target of a “true automated AI researcher” by March 2028. The Opus 4.6 and GPT-5.3-Codex system cards describe this tier as the present. ICML 2026 hosted an RSI workshop. Sergey Brin’s internal memo (via The Information and Fortune) directed DeepMind to make its models “primary developers” of code; Business Insider reported that Brin runs Gemini from a converted microkitchen at the DeepMind main office and that an internal program monitors some employees’ coding sessions to train Gemini’s coding capabilities. Roughly a thousand researchers are reportedly on related initiatives at DeepMind. Anthropic’s Institute has published a report on the tier (Dario Amodei, September 12, 2026; TIME, “Inside the Race to Make AI Build Itself,” August 7, 2026; Van Hollen letter; Fortune, September 3, 2026). Simon Lermen’s distinction between true RSI and AI-automated AI R&D is worth keeping; the tier one label follows his framing.
Tier two: component self-improvement. AI systems improving specific parts of their own training or inference stack with measurable gains. Google’s release note for Gemini 3.8 Flash and 3.8 Flash Cyber states that both models were “further accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models” (Google Blog, “Introducing Gemini 3.8 Flash and 3.8 Flash Cyber,” September 2026). Fortune’s account of Google’s May-to-August progression describes four Flash models shipped in 106 days: a two-agent self-improvement loop for a game in May, a three-agent loop helping train a robotics model in August, and loops refining the Gemini models themselves in September (3.8 Flash). AlphaEvolve, the DeepMind system for accelerating specific ML tasks, produced a matrix-multiplication kernel improvement of roughly twenty-three percent, which reduced Gemini’s training time by approximately one percent. The Navier-Stokes proof (ten thousand agents, eighty-eight hours, Lean-verified) and Neon (Liam Fedus’s 1,300 H200s over months of lab data, surpassing Astra on a materials benchmark) are evidence of what automated research produces, though neither is a self-accelerating loop.
Tier three: recursive self-improvement proper. A loop in which a system redesigns its own research process without human intervention and each cycle measurably accelerates the next. Nothing public meets the test. The skeptical check is that Google’s flagship Pro model has not shipped in months because internal prototypes showed insufficient progress over Flash to justify release; a lab that had achieved tier three would not have a stalled flagship. Evan Hubinger’s bet that his team would catch a self-improving misaligned model, and his statement that “there’s certainly a chance that we will just fail,” describe the alignment stakes at this tier. The September 2026 rumor (a leaker’s tweet whose capitalized letters spelled “RSI,” treated by parts of the industry as confirmation) is documented here only as evidence of the state of the discourse; Google employees are reported to be tempering expectations. Trending Topics’s test (several consecutive improvement cycles in which a system rebuilds its own research process without human intervention, at a measurably rising rate) is the specific standard against which nothing public currently qualifies (Trending Topics, “Three Letters Set the AI World Buzzing,” September 14, 2026).
The term “recursive self-improvement” migrated in roughly six months from AI-safety forums into a Google product release note, a Reuters report on a co-founder’s strategy, and a Business Insider profile. The labs now use it as a selling point. The AlphaEvolve figure (twenty-three percent for a specific kernel, one percent for training time overall) is the specific measurement that keeps the migration honest: it is real, it is concrete, and it is small. Tier three is not yet a demonstrated capability; tiers one and two are. The likelihood of the shade’s scenario moves from approximately twenty-five percent to approximately thirty-five percent on the basis that the precondition (tier one) is met and every major lab is pursuing tiers two and three explicitly. It does not rise further because tier three is undemonstrated and the measured loop magnitude is on the order of one percent per cycle at the component level.