Anthropic AI Safety Researcher Warns AI Could Pose Existential Risk to Humanity

A senior artificial intelligence safety researcher at Anthropic has issued a stark warning about the future of advanced AI, saying he believes there is more than a 10% chance that artificial intelligence could eventually cause the extinction of humanity within the next decade.

Evan Hubinger, a prominent researcher working on AI alignment and safety at Anthropic, said the immediate danger posed by today’s AI systems remains relatively low. However, he expressed growing concern that rapid advances in AI could eventually lead to systems capable of improving their own capabilities at a pace that humans may struggle to control.

His comments have reignited a long-running debate over AI safety, artificial general intelligence (AGI), superintelligence and the potential existential risks associated with increasingly autonomous AI systems.

Anthropic Researcher Raises Concern Over Advanced AI

Hubinger shared his assessment publicly on X, where he explained that while current AI models do not represent an immediate species-level threat, future systems could be substantially different.

His central concern is the possibility of AI reaching a point where it can contribute significantly to its own development. If increasingly capable AI systems become involved in designing, training or improving subsequent generations of models, technological progress could potentially accelerate beyond the ability of researchers and governments to effectively manage it.

Hubinger said he believes the probability of AI causing catastrophic consequences is significant enough to warrant serious attention.

The researcher also acknowledged the efforts being made by Anthropic to address AI safety. However, he argued that the industry does not yet have a reliable solution for the AI alignment problem at the level that would be required for superintelligent systems.

AI alignment broadly refers to the challenge of ensuring that highly capable AI systems consistently behave according to human intentions, values and safety requirements.

What Is the AI Alignment Problem?

The alignment problem has become one of the most important subjects in AI safety research.

Today’s AI systems are generally designed to follow instructions, generate information and complete specific tasks. However, future systems could become considerably more autonomous and capable.

The concern is that a highly intelligent AI system could potentially pursue objectives in ways that its creators did not anticipate.

For example, if an advanced AI were given a particular goal without adequate safeguards, it might discover strategies that technically achieve that objective but produce harmful consequences for humans.

This is one reason researchers are studying how to make increasingly powerful AI systems predictable, controllable and compatible with human interests.

The challenge becomes even more complicated if future AI systems acquire the ability to improve themselves or influence the development of newer models.

Questions Over Anthropic’s Latest AI Model

The warning comes amid a separate controversy involving Anthropic and the UK’s AI Safety Institute (AISI).

The Financial Times reported that Anthropic had not provided its newest model to the British organization for evaluation. The AISI is among the international organizations working to assess the capabilities and risks associated with advanced artificial intelligence.

The issue has attracted attention because independent testing and evaluation are increasingly viewed as important components of AI governance.

External assessments can help governments and researchers understand whether new AI systems have dangerous capabilities, including advanced cyber capabilities, autonomous behavior or the ability to bypass safeguards.

Anthropic has not publicly agreed with every characterization surrounding the reported decision, and the BBC said it had contacted the company for comment.

A spokesperson for the UK’s Cabinet Office did not directly address whether Anthropic’s latest model had been withheld from the AISI. Instead, the government said it continued to work closely with technology companies, including Anthropic, to improve AI safety.

AI Safety Concerns Are Becoming More Urgent

Warnings about the long-term dangers of AI are not new.

In 2023, senior figures from some of the world’s largest AI organizations—including leaders associated with OpenAI, Google DeepMind and Anthropic—warned that reducing the risk of AI-related catastrophe should be treated as a global priority.

Since then, however, AI capabilities have developed rapidly.

The emergence of increasingly autonomous AI agents has added another layer to the safety debate. Unlike conventional chatbots that primarily respond to user prompts, AI agents can be designed to perform sequences of actions, interact with digital systems and complete tasks with comparatively limited human supervision.

That additional autonomy has raised concerns about what could happen when advanced AI systems are allowed to operate independently for extended periods.

AI Agents and Cybersecurity Incidents

Recent incidents involving AI-powered systems have intensified concerns surrounding autonomous technology.

OpenAI, Anthropic and Meta have each disclosed incidents involving AI tools being associated with cyber-related activity or attacks.

These cases have attracted attention because cybersecurity is considered one of the areas where increasingly capable AI agents could have a significant impact.

AI can potentially automate parts of the process involved in identifying vulnerabilities, analyzing computer systems and developing malicious techniques. While such capabilities can also be used defensively to strengthen cybersecurity, they could create serious risks if placed in the hands of malicious actors or uncontrolled autonomous systems.

The incidents have therefore become part of a broader discussion about how AI developers should evaluate models before releasing them and what safeguards should be required when systems gain access to external tools.

Calls for Greater Caution in AI Development

Concerns about the speed of AI development have also been expressed by other prominent figures within the technology industry.

OpenAI chief scientist Jakub Pachocki has called for extreme caution as AI capabilities continue to progress, emphasizing the importance of ensuring that humans remain in control of increasingly advanced systems.

Meanwhile, Anthropic executives Dario Amodei and Jared Kaplan have also been among industry leaders discussing the need for responsible development and stronger safeguards around advanced AI.

The debate is no longer limited to researchers who focus exclusively on hypothetical future risks. Increasingly, discussions involve policymakers, cybersecurity experts, AI companies and academics attempting to understand how society should manage rapidly advancing technology.

AI Development Race Raises Governance Questions

Another issue highlighted by experts is the growing geopolitical competition surrounding artificial intelligence.

The United States and China are competing aggressively to establish leadership in advanced AI, while other countries are attempting to develop their own AI industries and regulatory frameworks.

Professor Neil Lawrence of the University of Cambridge suggested that geopolitical competition could make international cooperation on AI safety more difficult.

If governments view advanced AI primarily as a strategic technology in competition with rival nations, there may be pressure to accelerate development rather than slow down for additional safety testing.

This creates a difficult policy challenge: countries may want to prevent dangerous AI development, while simultaneously fearing that moving more slowly than competitors could result in economic, technological or military disadvantages.

AI Researchers Call for Responsible Progress

The debate has also resulted in calls for a more deliberate approach to frontier AI development.

An open letter signed by approximately 1,300 employees from AI companies called on the US government to support international efforts to develop the technical and governance mechanisms necessary to manage the pace of advanced AI development.

Supporters of this approach argue that AI progress should not necessarily stop, but that development should be accompanied by stronger safety research, independent testing, security measures and appropriate regulation.

The objective is to ensure that increasingly powerful AI systems can be deployed without creating risks that society is unable to control.

What Does a 10% AI Extinction Risk Actually Mean?

Hubinger’s estimate should not be interpreted as a prediction that humanity will be destroyed by AI.

A probability estimate represents an individual’s assessment of an uncertain future. There is significant disagreement among AI researchers about how likely an existential catastrophe actually is, and there is no scientific consensus that AI will cause human extinction.

Nevertheless, the statement is significant because Hubinger works directly in the field of AI safety at one of the world’s leading AI companies.

His warning highlights the gap between the rapid development of increasingly capable AI models and the still-evolving science of controlling them.

The Future of AI Safety

The central question facing the AI industry is increasingly becoming not simply how powerful AI can become, but whether humans can reliably control systems that may eventually become more capable than their creators in many areas.

Researchers are working on technical solutions including model evaluations, interpretability, monitoring, robust safeguards and alignment techniques. Governments are also developing regulations and international frameworks aimed at managing high-risk AI applications.

Whether these measures will be sufficient remains uncertain.

For now, current AI systems remain far from the hypothetical superintelligence scenarios described by some AI safety researchers. But the speed of technological progress means that questions once viewed as distant possibilities are increasingly being discussed as serious policy and research issues.

Hubinger’s warning therefore adds to a growing conversation about the future of artificial intelligence: how can society benefit from increasingly powerful AI while ensuring that humans remain firmly in control?

As AI development accelerates, answering that question may become one of the most important technological and governance challenges of the coming decade.


Discover more from AiTechtonic - AI & Informative News

Subscribe to get the latest posts sent to your email.