News and stories shaping America and the world

Search
×

Former OpenAI and Anthropic researcher warns about AI risks to humanity

Jacob Coxon left Anthropic and says AI companies are moving toward superintelligence without yet knowing how to guarantee that increasingly powerful systems will remain under human control

A warning from a researcher who worked at two of the world’s leading artificial intelligence companies is fueling a new debate over how far AI could go and, more importantly, whether humans will be able to keep it under control.

Jacob Coxon, 27, announced this week that he was leaving Anthropic, the company behind Claude. Before joining Anthropic, he also worked at OpenAI.

According to Coxon, he spent the past three years working directly on the pre-training of artificial intelligence systems at both companies.

When announcing his departure, the researcher raised a serious concern: in his view, neither OpenAI nor Anthropic is handling the rapid advancement of the technology responsibly enough.

Coxon says major AI labs are racing toward what is known as artificial superintelligence, or ASI, a hypothetical system that could outperform humans across virtually every intellectual task.

In one of his most widely shared statements, Coxon said people directly involved in developing these systems seriously believe artificial intelligence could pose a threat to human survival before the end of this decade.

He also accused the companies of “gambling with our lives.”

Anthropic researcher reinforces the warning

The story gained even more attention after Evan Hubinger, an Anthropic researcher who works on AI alignment and safety, publicly responded to Coxon’s statements.

Hubinger said the concern is real and estimated there is a greater than 10% chance that extremely advanced artificial intelligence could lead to human extinction within the next decade.

That figure is not a scientifically established probability or a confirmed prediction. It represents Hubinger’s personal assessment of the risk.

Hubinger also made an important distinction: according to him, the risk posed by AI models available today is low. His concern is focused on what could happen if artificial superintelligence is developed and becomes capable of rapidly improving itself.

The researcher also acknowledged that Anthropic is working to reduce these risks but said the company does not yet have a definitive solution to the alignment problem that could emerge with future superintelligent systems.

but how could AI pose such a risk?

This is where some of the claims circulating on social media need additional context.

Coxon did not reveal the existence of a secret program designed to “train artificial intelligence to kill humans.”

The warning involves a different issue.

Researchers are concerned that much more powerful and autonomous AI systems could be given goals by humans and, while pursuing those goals, develop unexpected strategies or behaviors that become difficult or impossible to control.

Scenarios discussed by AI safety researchers include the use of artificial intelligence to help develop biological weapons, conduct large-scale cyberattacks, compromise critical infrastructure or operate digital systems with increasing levels of autonomy.

There is also a broader concern known as recursive self-improvement.

Under this scenario, an artificial intelligence system could help develop an even more powerful version of itself. That new system could repeat the process, potentially accelerating technological development far beyond the pace of human research.

The concern is that humans could eventually lose the ability to fully understand or control the decisions made by these systems.

do today’s AI systems pose this danger?

There is no evidence that ChatGPT, Claude or other AI models currently available to the public are on the verge of causing human extinction.

Anthropic itself published an assessment of sabotage risks involving its models and concluded that current systems pose a very low, though not entirely nonexistent, risk of carrying out autonomous, misaligned actions that could significantly contribute to catastrophic outcomes.

The company also said it has moderate confidence that Claude Opus 4 does not have consistently dangerous goals or sufficient capabilities to carry out complex sabotage strategies while avoiding detection.

In other words, the debate raised by these researchers is primarily focused on AI systems that have not yet been developed.

CNN covers the growing concerns

CNN devoted airtime to the warning and the broader debate surrounding advanced artificial intelligence.

In one discussion, CNN artificial intelligence correspondent Hadas Gold explained that these concerns are not new. Executives and researchers within the AI industry have been discussing for years the possibility that highly advanced systems could eventually become difficult to control.

What made Coxon’s case particularly significant was his direct involvement in training models at both OpenAI and Anthropic, followed by his decision to publicly leave the company while raising these concerns.

CNN also highlighted comments from another Anthropic employee, Samuel Marks, who said people working directly on artificial intelligence consider it possible that the technology could have extremely serious consequences and that, in general, more experienced researchers within the industry tend to express greater concern.

why is this debate happening now?

The discussion is gaining momentum because AI models are evolving beyond systems that simply answer questions.

So-called AI agents can already perform sequences of tasks, write and execute code, use digital tools, research information and make certain decisions with less direct human involvement.

For researchers focused on AI safety, the more autonomous these systems become, the more important it is to understand, monitor and limit what they are capable of doing.

Coxon argues that AI development is moving too quickly for those safeguards to be created at the same pace.

what Anthropic says

Anthropic conducts dedicated research into AI alignment and safety and has publicly acknowledged potential risks associated with increasingly powerful artificial intelligence systems.

In a report examining sabotage risks, the company said its current models have not yet reached the capability level that would trigger some of the additional safeguards designed for more advanced systems.

Anthropic has also developed a framework called its Responsible Scaling Policy, which is intended to introduce stronger safety measures as AI models become more capable.

Hubinger himself, despite raising concerns about the potential risks, has said he believes Anthropic is trying to address the problem.

The central question in the debate is whether the AI industry will be able to develop effective safeguards before artificial intelligence systems reach levels of capability that researchers themselves do not yet know how to fully control.

Sources: CNN, CBS News, Anthropic, TechCrunch, and public statements from Jacob Coxon and Evan Hubinger.

Picture of LFW Editorial Team

LFW Editorial Team

The LFW Editorial Team produces and curates news, stories and original content for LFW Portal, connecting people, communities and perspectives across the United States and around the world.

You might be interested in this :

Leave a Reply

Your email address will not be published. Required fields are marked *