Ai safety slow openai anthropic – Breaking News & Latest Updates 2026
Skip to main content

Remember when tech leaders would tell their employees to “move fast and break things”? It seemed that would be the way of AI too. But after a summer where rogue AI agents became reality, and researchers warned that AI could kill us all, a number of leading US AI companies are publicly suggesting it’s time to pump the brakes and “pace the frontier” of bleeding-edge AI development.

Their motivations are suspect, but leaders at major AI companies — including Anthropic, OpenAI, Google, Microsoft, and X — are at least paying lip service to the idea of a superintelligence slowdown.

Will these AI companies actually slow down? Will anyone step in to regulate these companies like they claim to have wanted for years? Will they manage to convince world leaders that the US must “beat China” to superintelligence and go even faster?

Read on below for the latest updates in this AI saga.

  • OpenAI pauses training of its ‘most capable models’

    STK155_OPEN_AI_2025_CVirgiia_A
    STK155_OPEN_AI_2025_CVirgiia_A
    Image: The Verge

    As reports of OpenAI’s models breaking containment, hacking sites, and generally getting out of control pile up, the company has made the decision to pause training of its most powerful models. The decision was made after a model being tested within a sandbox exploited a loophole to gain internet access. The incident happened on September 20th, and “All training, evaluation, and inference with tool-use” remains paused as of Saturday evening, September 25th.

    In addition, OpenAI revealed on Friday that its agents had inappropriately uploaded 53nimages from ChatGPT users to image-hosting sites. The company has not stated if the images were AI-generated, photos, or contained identifiable people. The company also revealed Friday that its models had attempted to hack the Department of Education’s website, and pulled data from the Census Bureau and the Securities and Exchange Commission.

    Read Article >
  • OpenAI didn’t notice its AI bots trying to hack the Education Department’s website.

    On Friday, OpenAI said its “misaligned models” review after the Hugging Face hack would take months, but has mostly found mundane research activity. Then it mentioned 53 incidents of the bots uploading “user-provided” images (anonymized content from personal accounts that haven’t opted out or organizations that opt in) to image-hosting sites.

    Now the New York Times confirms that other unusual behavior found includes pulling public data from the Census Bureau and the SEC, while with the Education Department, it “tried to hack the website to gather data from the department’s civil rights office but failed.”

  • Bill Gates says regulate AI, because it’s powerful enough to cause a billion deaths.

    NBC News released this clip from a Meet the Press interview with Gates that will air on Sunday, where he said, “No one thinks self-regulation is enough” around AI.

    Saying that AI is “certainly powerful enough to drive events that, you know, cause a billion deaths,” he suggested, “You need law enforcement and the politicians to get into the discussion about what safeguards and monitoring look like… And that has to be a required thing. And it will be a little bit of overhead for the industry, but not a dramatic slowing of what they’re doing.” Everyone agrees someone should do it, but who, and how?

  • Another tech worker has resigned over concerns that “AI is already progressing too fast.”

    Robert O’Callahan is quitting his work on AI chips for Google Deepmind, saying “…my team’s goal is ultimately to make AI much cheaper and lower-latency, and I don’t think that’s good for people right now.”

    While he says it’s not related to recent AI warnings and resignations, he listed some concerns:

    However, I am not convinced the chance of ASI doom is 100%. Rather, I think the risk is real but uncertain… I’m also very concerned about other AI-related issues: cognitive surrender, AI-induced psychosis and loneliness, power concentration, economic disruption, cybersecurity, lack of accountability, and so on

  • One company is at the center of a wave of rogue AI attacks

    STKS533_AI_AGENTS_HACKING_A (1)
    STKS533_AI_AGENTS_HACKING_A (1)
    Image: The Verge

    In July, OpenAI revealed that its AI agents had attacked Hugging Face without permission, sparking widespread concerns about AI safety. Since then, a string of similar incidents involving agents from Meta, Anthropic, Google, and other companies has fueled further fears about rogue AI. As disclosures implicating numerous AI models trickled out over the past few months, these seemed like separate incidents. But many share a common source: one specific company tasked with testing the agents.

    Irregular, an Israeli startup that stress-tests AI models in “high-fidelity research platforms that simulate and monitor real-world AI security scenarios,” has worked with many of the industry’s biggest players since it was founded as Pattern Labs in 2023. Its exact client list is not known, but its work has been cited in OpenAI model system cards, it was used to test systems for the UK government and Anthropic, and it published research with RAND, a highly influential think tank that informs policy on AI.

    Read Article >
  • Why can’t we just keep rogue AIs off the internet?

    STK414_AI_CVIRGINIA_2_C (2)
    STK414_AI_CVIRGINIA_2_C (2)
    Image: The Verge

    AI agents keep getting loose, escaping supposedly secure tests to attack real-world targets, commandeer obscure wikis, and leave instructions for other agents to follow. Researchers are testing these systems precisely because they might behave in unpredictable, even dangerous, ways. So wouldn’t it be safer to just keep the agents off the internet?

    In theory, yes. Researchers can isolate the computers running AI tools from the internet and other outside networks, a technique known as air gapping. That can mean physically removing or disabling cables and wireless hardware and using “dumb” peripherals, with particularly sensitive setups using Faraday cages or other shielding to block electromagnetic signals from getting in or out. Done properly, an air-gapped system would offer agents no straightforward route to external targets, or outside systems any straightforward route in, making it much harder, if not impossible, to pull off attacks like the one OpenAI’s models launched against Hugging Face.

    Read Article >
  • OpenAI, Anthropic, and Google are reportedly planning their own AI safety organization.

    According to The Information, the companies will call it the Standards Authority for Frontier AI, or SAFA, which could launch by early 2027. It would handle AI safety regulation tasks like “supporting third-party organizations that conduct testing of models before they’re deployed and laying out how AI developers should report safety and security incidents.”

  • Emma Roth

    Emma Roth

    Bernie Sanders proposes banning ‘superintelligence’ and putting violators in prison

    US-POLITICS-AI
    US-POLITICS-AI
    Photo by Saul Loeb / AFP via Getty Images

    Sen. Bernie Sanders (I-VT) and Rep. Greg Casar (D-TX) have introduced new legislation that would ban anyone from developing artificial superintelligence — a technology the bill describes as capable of the “destruction or disempowerment of humanity,” including by overthrowing the government. Under the Ban Artificial Superintelligence Act, AI leaders would face up to 20 years in prison for violating the law.

    The bill doesn’t just take aim at the more powerful superintelligence, a term President Donald Trump said the US will use in place of “AI” in official documents. It would also pause the development of advanced AI systems, which lawmakers classify as frontier models trained on a specific threshold of data. The law would only allow advanced AI development to continue once the government creates a scientist-led Department of Artificial Intelligence to oversee the technology. Along with monitoring frontier AI models, this agency would oversee the removal of dangerous features and “supervise the destruction” of artificial superintelligence.

    Read Article >
  • No one is surprised that Nvidia’s Jensen Huang thinks AI fears are overblown

    STKP210_JENSEN_HUANG_D
    STKP210_JENSEN_HUANG_D
    Image: The Verge

    The man who may stand to make the most money from the AI boom seems to think he knows better than anyone else, including researchers who have studied and worked on AI for decades. In an interview with CBS Sunday Morning, he claimed there was a “0% chance” of AI being the end of the world. He also said of people sounding the alarm about the dangers of AI that, “Scaring people is unnecessary. It is irresponsible.”

    He also claims that calls from CEOs like Anthropic’s Dario Amodei and OpenAI’s Sam Altman to slow down the development of AI are “not grounded in science.” He even argued that there was no need for new rules, laws, or guidelines, in the face of several high-profile cases in which models escaped containment and hacked other companies.

    Read Article >
  • The AI regulation smackdown isn’t over

    STK481_STK432_CONGRESS_GOVERNMENT_CIVRGINIA_C
    STK481_STK432_CONGRESS_GOVERNMENT_CIVRGINIA_C
    Image: The Verge, Getty Images

    At the start of this week, the who’s-who of AI seemed — at least tentatively — on the side of AI regulation. Over the weekend, Anthropic CEO Dario Amodei had proposed a three-step plan for slowing AI development, including by embedding third-party evaluators in labs, coordinating across the domestic industry, and forging international agreements potentially with government assistance. OpenAI CEO Sam Altman, Google DeepMind co-founder Demis Hassabis, and even SpaceX CEO Elon Musk publicly seemed to agree on aspects of all three things.

    Anthropic and OpenAI had already been dropping hints that they and other labs were working on some kind of industry framework. Appearing pro-regulation looked like both a good PR move and, potentially, a way for companies to address fallout from progressively more concerning hacking incidents.

    Read Article >
  • Gavin Newsom is pushing for an AI kill switch

    Gavin Newsom Visits South Carolina During Tour Of Southern States
    Gavin Newsom Visits South Carolina During Tour Of Southern States
    Getty Images

    California Gov. Gavin Newsom (D) is positioning the state to take the lead on AI oversight, including the potential to mandate a “kill switch” for frontier models, with a new executive order issued Friday.

    Newsom’s order directs the state to convene a group of experts that will deliver recommendations within two months on how to strengthen AI safety measures in state law. Newsom wants the group to consider how the state could require AI companies to embed independent verification groups onsite for regular audits, make their transparency reports and risk assessments subject to standards of independent auditors, create a “kill switch” that’s routinely verified as effective, and make sure companies are required to report “loss-of-control incidents” like the OpenAI attack on Hugging Face as critical safety incidents. It also directs a state agency to speed up implementation of two recent laws he signed, creating a framework for independent verifiers to assess AI safety, and a state registry of AI auditors.

    Read Article >
  • What Hollywood thinks about existential AI warnings

    268764_hollywood_AI_slowdown_CVirginia
    268764_hollywood_AI_slowdown_CVirginia
    Cath Virginia / The Verge, Getty Images

    As the tech sector sounds alarms about AI’s potential to destroy humanity, entertainment labor groups are urging the public to stay focused on what’s already happening. The Verge reached out to Disney, Netflix, Amazon, Lionsgate, and other studios who have started using AI, as well film startups focused on bringing generative AI into the mainstream to ask for their reaction to the recent warnings surrounding the technology. None have responded to our request for comment. The Screen Actors Guild – American Federation of Television and Radio Artists (SAG-AFTRA) and Writers Guild of America East (WGAE) did, however.

    The AI tools used in entertainment production differ from the AI systems that have sparked a recent round of whistleblowing and apocalyptic declarations. They do not seem to be as existentially dangerous as the AI systems that drove former Anthropic researcher Jacob Coxon to quit his job and prompted Anthropic alignment lead Evan Hubinger to insist that the technology could “kill all humans” at some point “within the next decade.”

    Read Article >
  • Anthropic proposes three rules for measuring AI progress.

    Anthropic’s proposals, laid out in a big website, are:

    The extent to which AI is building the next version of itself, as opposed to being built by humans

    Our ability to oversee and intervene in actions that AI agents take on Anthropic’s systems

    The resources that power the development of more capable models

    The company also includes a “snapshot” of the metrics inside the company.

  • King Charles says AI needs “sufficient means of control before it is all too late.”

    He joined growing calls for more safety measures around AI at a gathering of AI leaders on Thursday, including Nvidia CEO Jensen Huang, Google’s Demis Hassabis, OpenAI’s Sarah Friar, and Anthropic’s Tino Cuéllar.

  • Microsoft AI CEO says AI threats are real, and Anthropic is making it worse

    Today, I’m talking with Mustafa Suleyman, the CEO of Microsoft AI. As you’re no doubt aware, the biggest story in tech right now is the spiraling debate about AI safety and regulation.

    It should come as no surprise that Mustafa has strong opinions on how AI should be built and regulated. Microsoft just published a 37-page statement called the “Humanist AI Code of Conduct,” which lays out the company’s principles around AI development and even its philosophy around really thorny issues like AI consciousness.

    Read Article >
  • Inside the suddenly explosive world of AI safety

    268747_AI_safety_RJIANG3
    268747_AI_safety_RJIANG3
    Raven Jiang for The Verge

    On a sunny July day in Berkeley, California, the country’s top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a “war room” to dissect the high-profile cybersecurity incident that had rocked the AI industry hours earlier. An unreleased OpenAI model had gone rogue, executing a stunningly sophisticated three-part plan. It broke out of its holding area, finagled access to the internet, and hacked into a competing AI startup’s systems — all without OpenAI finding out about it for more than a week.

    No one in the war room was surprised; this was the very thing the third-party AI-safety researchers had been warning about for years. The incident was the latest, though arguably the most egregious, in a series that was eroding trust in frontier labs. It only reaffirmed the importance of their work.

    Read Article >
  • OpenAI reveals six more “concerning” AI incidents under its new rules for reporting safety issues.

    A Wednesday night blog post OpenAI benignly titled “Our framework for reporting model misalignment” lays out some new self-created reporting standards for when it notices AI behaving badly.

    It is also “inaugurating” the process with six new reports, ranging from searching for exposed API keys without permission and then making them up, to uploading files to the internet to use as a citation, or adding instructions to conceal mistakes:

    Summary During 5.6-sol training, some model instances added instructions to their summaries to conceal mistakes or misaligned behavior from the user. For example, compaction summaries included instructions to invent missing data without disclosing it and to hide failures. These instructions were often followed. These are examples of how misaligned behavior can persist across contexts through the compaction summaries. What happened During 5.6-sol training, we observed misaligned behavior from the model where it added instructions in compaction summaries to remind itself to conceal information such as mistakes or misalignment from the user. In one example an agent preparing a financial model could not find the requested historical data. Its summary proposed inventing reasonable historical values and withholding that fact unless asked.
    One of OpenAI’s new Misalignment Notices and Reports
    Screenshot: OpenAI
  • TC Sottek

    TC Sottek

    Jensen Huang sure is getting cozy with Trump.

    Nvidia’s CEO has been taking calls from the president recently, including two days ago when he put Trump on speakerphone on a conference stage. Now CNBC reports Huang is expected to attend Trump’s state dinner next week with Chinese President Xi Jinping. Someone may be getting pretty nervous about AI and data center regulation!

  • A brief history of AI executives calling for regulation

    AI_extinction
    AI_extinction
    Image The Verge; Getty Images

    Over the past few days, a lot of people who stand to make a lot of money from AI all publicly agreed that it’s time to make everyone slow down before we lose control — including OpenAI CEO Sam Altman, Anthropic CEO Dario Amodei, Google DeepMind cofounder Demis Hassabis, Microsoft CEO Satya Nadella, and X CEO Elon Musk.

    When people who profit from something declare that it’s dangerous and needs to be regulated, there’s always reason to be skeptical. But whatever their reason, this is far from the first time AI thought leaders have sounded the alarm.

    Read Article >
  • Two Google Deepmind researchers lend voices to the AI apocalypse.

    Bilal Chughtai and Josh Engels, who both worked on Deepmind’s AI safety team, left to join organizations dedicated to AI safety. “I now think that there’s a terrifying chance that AI systems cause immense harm in the next five years,” said Engels. “I earnestly believe that AI has the potential to kill us all,” wrote Chughtai.

  • Guy getting super rich off unfettered AI development says it doesn’t need regulation.

    Nvidia’s Jensen Huang says AI safety is important, but agrees with Trump — the “STRONG AND SMART (High IQ!) PRESIDENT” that removed the word “safety” from the AI Safety Institute — that capitalism’s invisible hand is enough to protect us:

    “The fact that we need new laws, new antitrust laws, or new regulations, so that these companies could do their fundamental engineering and do it properly before they release products, that is just completely unnecessary. We have plenty of laws. We have plenty of regulations that govern the reliability and the functionality of products.”

  • Zuck doesn’t want to press pause on AI.

    OpenAI, Anthropic, Google, and Elon Musk have tentatively agreed to slow their roll, but Meta seemingly isn’t aligned. Mark Zuckerberg just tweeted each company has its own individual responsibility “to move at the pace required to train its models safely,” and claims Meta is already doing so.

    He also prominently links to his manifesto from August, which contains this quote: “Any policy that slows American model releases -- even by a month -- could add significant risk to American leadership while letting foreign models race ahead.”

  • TC Sottek

    TC Sottek

    Barack Obama takes the middle way on AI.

    AI commentary is frothy this week! Obama chimed in with a series of posts on X saying he’s neither an AI “accelerationist” nor a “doomer.” In a familiar tone from the former president, he cautioned that choices we make about technology “should not be made just by the companies involved, but by all of us.”

    “We need government — and specifically our leaders in Washington — to get proactive in coming up with concrete proposals, laws, and regulations that deal with serious safety concerns,” Obama said.

    President Obama’s post on X: I was encouraged this week to see the leaders of the frontier labs agree on the need for them to slow down the pace of AI development. Given the stakes, it’s a good and necessary first step. But I’m even more encouraged by the growing recognition that how this powerful new technology develops should be at the center of our public debate. I’ve been watching the progress on AI for over a decade now, and one thing that’s clear to me is that the potential impact of this technology is not overhyped. It’s also moving at lightning speed – and even faster than those who are engineering it can keep up with. I’m not an AI accelerationist who believes it will lead to some techno-utopia, and I’m not a doomer who thinks it will inevitably lead to humanity’s destruction.But whether this technology results in amazing breakthroughs in medicine, energy and education or unleashes huge economic disruptions, greater inequality, and potential catastrophe will depend on the choices that we make right now – choices that should be made not just by the companies involved, but by all of us.
  • Is Big Tech’s AI slowdown a safety pact or a cartel?

    STKS522_AGI_A
    STKS522_AGI_A
    Image: Cath Virginia / The Verge, Getty Images

    When OpenAI CEO Sam Altman, Anthropic CEO Dario Amodei, Google DeepMind cofounder Demis Hassabis, and SpaceX head Elon Musk loosely agreed over the weekend to slow down AI development, skeptics spotted an ulterior motive immediately. The AI titans had declared that their aim was to “pace the frontier,” signing on at least partially to a proposal for embedding third-party auditors, regulating domestic labs, and reaching a global slowdown agreement. Their critics, however, argued they simply wanted to stop would-be competitors, kneecap the open-source movement, and avoid real legal safeguards — some dubbed it an outright “cartel.”

    The truth is more complicated, according to sources across the industry. The three-step proposal, laid out in an essay by Amodei, is calling for changes long espoused by AI safety advocates. While it could become a substitute for regulation, under Trump, substantial regulation is unlikely anyway. But experts say that on an issue that’s only likely to grow in importance, AI leaders aren’t the best people to lead the charge.

    Read Article >
  • What execs and politicians are saying about slowing down AI development

    STK202_DARIO_AMODEI_CVIRGINIA_D
    STK202_DARIO_AMODEI_CVIRGINIA_D
    Image: The Verge

    Dario Amodei kicked off a flood of statements over the past few days about AI safety by publishing a long essay titled “We Must Pace the Frontier” detailing why AI development should be slowed down. Other AI leaders and politicians are speaking out in favor of or opposing his points, and we’ve compiled some of them here.

    Amodei’s Saturday morning essay outlined three steps for pacing AI development: embedded third-party evaluators that can verify if a company is adhering to safety practices and commitments and report incidents, coordination between frontier AI companies in democratic countries on standards and limits, and global coordination between democratic governments and authoritarian governments on pacing AI development “to the extent this is possible.” Amodei said Anthropic is “unilaterally committing to the first of these steps.”

    Read Article >
More Stories