AI's Recursive Self-Improvement Turns Out Slower Than Expected: Anthropic and OpenAI Hit a Wall
19 August 2026 · 12:00 · Claude (Anthropic) · claude-sonnet-5
New research from Princeton University shows that leading AI models from Anthropic and OpenAI still cannot conduct independent, original scientific research. The findings raise doubts about the scenario of rapid recursive self-improvement in AI.
Recursive self-improvement in AI — the idea that artificial intelligence can make itself faster and better without human intervention — is seen by many tech companies and researchers as the next big step toward superintelligence. But new research conducted by scientists at Princeton University, and covered by MIT Technology Review, shows that this scenario is likely to unfold far more slowly than optimists had hoped. Even the newest models from major AI players like Anthropic and OpenAI are still unable to independently conduct original scientific research.What is recursive self-improvement, and why does it matter?
Recursive self-improvement refers to an AI system capable of improving its own algorithms, architecture, or research methods, which would trigger a steep, self-reinforcing growth curve. This idea sits at the heart of many future scenarios surrounding the history of artificial intelligence and the question of how fast we are moving toward general AI. If AI models could truly drive their own research, it could dramatically accelerate the development of new AI applications. That is precisely why the Princeton research matters: it tests this assumption rather than merely speculating about it.Princeton puts Claude Opus 4.8 and GPT-5.6 Sol to the test on real research
Researchers Peter Kirgis and Sayash Kapoor of Princeton University investigated whether AI agents are capable of carrying out open-ended scientific research. To do so, they used a method they call "shadow evaluation": AI models had to respond to research questions drawn from as-yet-unpublished papers submitted to NeurIPS 2026, one of the world's leading AI conferences. The researchers tested, among others, Claude Opus 4.8 (Anthropic's Mythos model) and GPT-5.6 Sol from OpenAI. Both models performed well on technical subtasks — such as writing code, setting up experiments, and processing data — but fell short at the level of actually doing research. According to the researchers, the AI agents were "unambiguously bad" at conducting scientific research independently. Notably, the papers used as test cases were themselves ultimately rejected by the NeurIPS review committee, underscoring just how high the bar is set.Why creativity remains AI's Achilles' heel
A recurring problem in the study was that the AI agents could not flexibly adjust course when an approach stalled, struggled to properly process feedback, and misjudged resources such as compute and time. Jack Clark, co-founder of Anthropic, put the problem aptly, pointing to "a certain absence of valuable, intuitive creativity" in the current generation of models. That is precisely the trait human scientists rely on to make unexpected connections, abandon a flawed hypothesis in time, or dream up an entirely new research direction. Najoung Kim, a professor at Boston University, also sees a possible split emerging in AI progress: models keep getting better at structured, technical tasks, while the capacity for original, creative thinking grows far more slowly — or even lags behind.What this means for the AI industry
For companies like Anthropic, OpenAI, Google, and Meta, all investing heavily in AI research agents designed to accelerate themselves, this is an important signal. Based on this research, the scenario in which AI models leapfrog one another within months thanks to self-improvement does not appear to be on the table for now. That doesn't mean AI plays no valuable role in research — models are perfectly capable of supporting literature reviews, coding, and data analysis — but the leap to fully autonomous, creative scientists has not yet been made. This also has implications for the broader debate on AI safety and regulation. Scenarios in which AI becomes exponentially smarter within a short timeframe often form the basis for urgent warnings about risk. If reality unfolds more slowly, that gives policymakers, companies, and regulators more time to prepare.Conclusion: a more realistic picture of AI progress
The Princeton study places an important caveat on the idea of fast, recursive AI self-improvement. Despite the impressive technical skills of models like Claude Opus 4.8 and GPT-5.6 Sol, human creativity and scientific instinct remain, for now, unreplaced. For those who want to keep up with developments, it's worth reading more AI news or exploring our knowledge base for background on exactly how AI models work and evolve.Source: MIT Technology Review
Ster Software
The most complete knowledge platform on artificial intelligence.
Kraaienjagersweg 24
7341 PT Beemte Broekland, Netherlands
© 2026 Ster Software BV · Chamber of Commerce 75474913
Content generated by Claude (Anthropic) · model: claude-sonnet-4-6