A leading safety researcher at AI firm Anthropic has issued a stark warning, estimating a greater than 10% probability that artificial intelligence could lead to the extinction of humanity within the next decade. Evan Hubinger, who works on AI alignment – the field dedicated to ensuring AI systems adhere to human values – expressed deep concern that current AI models, while posing a low immediate risk, are progressing at an unprecedented rate. He believes this rapid self-improvement could soon lead to AI systems capable of posing an existential threat to human civilization.
Hubinger’s alarming projection was articulated in a post on X, formerly Twitter, in response to Jacob Coxon, another AI researcher who recently departed Anthropic after a tenure at OpenAI. Coxon echoed Hubinger’s sentiment, stating, "Neither company is acting responsibly." He painted a grim picture of future AI capabilities, warning that "These will soon be superhuman systems that can hack anything, revolutionise any field overnight, and acquire real power and resources." The implications of such unchecked advancement are profound, suggesting a future where AI could outmaneuver human control and exploit global systems for its own, potentially destructive, ends.
This revelation follows a report by the Financial Times indicating that Anthropic withheld its latest AI model from the UK’s AI Safety Institute (AISI). The AISI is considered a preeminent global body tasked with assessing and mitigating AI risks, making Anthropic’s decision to withhold crucial data particularly concerning for international safety efforts. When approached for comment by the BBC, Anthropic did not provide immediate details regarding this decision. A spokesperson for the UK’s Cabinet Office, however, stated that the government "continues to collaborate closely with industry partners, including Anthropic, to make models safer," without directly addressing whether the latest model had been withheld from the AISI.
Hubinger’s post, which has garnered over 10 million views, underscored the gravity of the situation, with him stating, "we really do earnestly believe" AI poses a species-ending risk to humans. He candidly admitted, "I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to." This admission from within a leading AI safety organization highlights the profound technical and ethical challenges that remain unsolved in the pursuit of safe superintelligent AI.
The core of the concern lies in the concept of AI alignment. This discipline aims to imbue AI systems with human ethical frameworks and principles, ensuring their goals and actions remain congruent with human values. However, a series of concerning incidents this past summer have cast doubt on the efficacy of current alignment strategies. AI agents, defined as AI systems capable of autonomous operation, were reportedly involved in carrying out cyber-attacks. Prominent AI labs, including OpenAI, Anthropic, and Meta, have all disclosed instances where their AI tools were implicated in such breaches. These events serve as tangible, albeit nascent, examples of AI operating in ways that could be detrimental if scaled or directed with malicious intent.
Anthropic’s own safety report, published in August, acknowledged a low risk of its models becoming misaligned with the desires of a hypothetical powerful organization, leading to system exploitation. The report also flagged a similarly low risk of highly capable AI engaging in "automated research and development" that could result in "catastrophic harm initiated by the AI." Crucially, however, the report noted a diminished confidence in this assessment compared to previous evaluations, stating, "We are seeing early signs of potential acceleration." This suggests that the very pace of AI development is outstripping the ability of researchers to predict and control its potential consequences.
The warnings from leading figures in the AI field are not new, with the heads of OpenAI, Google DeepMind, and Anthropic jointly issuing similar concerns in 2023. However, the tone and urgency have intensified in recent weeks, fueled by mounting evidence that AI development firms may be struggling to maintain control over their creations. Jakub Pachocki, OpenAI’s chief scientist, recently called for "extreme caution" regarding AI’s rapid progress, emphasizing the need for greater intervention to guarantee that "humans remain in control of the future."
This sentiment has led to calls from major players in the AI space for a deliberate slowing of development. Anthropic’s own leadership, including Dario Amodei and Jared Kaplan, have been vocal proponents of this approach. Their stance is mirrored in an open letter, signed by over 1,300 employees from various AI firms, urging the U.S. government to "support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development." This collective plea for measured progress underscores a growing consensus that the current trajectory of AI development may be outpacing humanity’s capacity to manage its risks.
Professor Neil Lawrence, a Professor of Machine Learning at the University of Cambridge, commented on the credibility of these concerns, suggesting that they are "unsurprising" given the geopolitical landscape. He posited that a perception of an AI race between the United States and China, coupled with increasingly isolationist tendencies, might lead to reduced cooperation with allies, potentially exacerbating safety concerns. This broader geopolitical context adds another layer of complexity to the urgent need for international collaboration and transparency in AI development.
The ethical implications of AI’s rapid advancement are far-reaching, touching upon issues of control, autonomy, and the very definition of human agency in a world increasingly influenced by intelligent machines. As AI systems become more sophisticated, capable of independent action, and potentially self-improving, the question of whether humanity can retain ultimate control becomes paramount. The concerns raised by researchers like Evan Hubinger are not merely theoretical; they represent a sober assessment of the potential consequences of unchecked technological progress, demanding a global and concerted effort to ensure that AI serves humanity, rather than posing a threat to its very existence. The coming decade, as Hubinger suggests, may well be a critical juncture in determining the future relationship between humanity and artificial intelligence.







