A cluster of public resignations from AI safety teams in September 2026 has pushed 'superalignment' - the problem of keeping AI systems far more capable than humans steerable by human intent - back into mainstream policy debate.
Jacob Coxon, a pretraining researcher who had worked at both OpenAI and Anthropic, wrote on 9 September 2026 that neither company was acting responsibly and that the labs were 'racing straight to self-improving superintelligence'.
On 11 September 2026 Joe Benton, who led Anthropic's Scalable Oversight team, said he had left the company two weeks earlier and was moving to METR to work on autonomous-AI risk evaluations.
Superalignment differs from ordinary alignment because human supervisors can no longer audit the reasoning of a system smarter than themselves - the oversight asymmetry problem.
OpenAI set up a Superalignment team in July 2023 with a pledge of 20% of its secured compute and dissolved it in May 2024 after both co-leads, Ilya Sutskever and Jan Leike, resigned.
Ordinary alignment assumes a human can read an AI's answer and judge whether it is right. Superalignment asks what happens when the system's reasoning exceeds human comprehension, so the supervisor can no longer verify what is being approved.
Simple Analogy: A non-player asked to sign off on a grandmaster's chess move they cannot evaluate.
| Aspect | Alignment (narrow) | Superalignment |
|---|---|---|
| Capability assumed | AI at or below expert human level | AI far above human level |
| Who supervises | Human reviewers, directly | Weaker aligned models and automated evaluators |
| Main method | Human feedback on outputs | Weak-to-strong generalisation, automated oversight, interpretability |
| Central difficulty | Specifying what humans want | Verifying reasoning humans cannot follow |
| Failure feared | Biased or harmful outputs | Loss of control over a self-improving system |
OpenAI announces a Superalignment team co-led by Ilya Sutskever and Jan Leike, pledging 20% of secured compute over four years.
Sutskever and Leike leave OpenAI; the Superalignment team is dissolved and its members reassigned. Leike subsequently joins Anthropic.
India hosts the AI Impact Summit at Bharat Mandapam, New Delhi; the New Delhi Declaration on AI Impact is adopted and the IndiaAI Safety Institute is announced.
Jacob Coxon publishes his resignation posts, accusing OpenAI and Anthropic of 'gambling with our lives'.
Joe Benton reveals he has left Anthropic's Scalable Oversight team and will join METR for autonomous-AI risk evaluation.
US AI laboratory whose Alignment Science and Scalable Oversight teams are at the centre of the September 2026 resignations
Created the Superalignment team in July 2023 and dissolved it in May 2024; Coxon worked on pretraining here before joining Anthropic
Non-profit that evaluates frontier AI models for dangerous autonomous capability; Joe Benton's new employer
India's national body for indigenous, science-based research on AI safety, risk assessment and governance relevant to developing countries
Build India's artificial-intelligence ecosystem across compute, datasets, applications, skilling, startup finance and safety
Key: Approved by the Union Cabinet in March 2024 with an outlay of Rs 10,371.92 crore over five years, under MeitY; the IndiaAI Safety Institute sits in its 'Safe & Trusted AI' pillar
Set a shared global agenda for AI that is inclusive, trustworthy and development-oriented
Key: Adopted at the India AI Impact Summit, New Delhi, February 2026; endorsed by 92 countries and international organisations and structured around seven 'Chakras' or pillars
GS Paper III > Science and Technology - developments and applications in IT and artificial intelligence; GS Paper II > Governance and regulation of emerging technologies
General Awareness > Science and technology in current affairs
With the present state of development, Artificial Intelligence can effectively do which of the following? 1. Bring down electricity consumption in industrial units 2. Create meaningful short stories and songs 3. Disease diagnosis 4. Text-to-Speech Conversion 5. Wireless transmission of electrical energy Select the correct answer using the code given below:
Answer: 1, 3 and 4 only
AI safety research aimed at keeping systems far more capable than humans reliably under human control, in conditions where direct human verification is no longer possible.
The tendency of a goal-directed system to develop unintended sub-goals such as self-preservation or resource acquisition while pursuing any assigned objective.
Examining the internal workings of a neural network to understand and verify how a given behaviour is produced.