AI: From Safety Warnings to the White House Agreement
Explore the AI safety warnings, employee concerns, and major incidents that preceded the September 29, 2026 voluntary agreement at the White House.

Over time, AI models have improved and become part of everyday life. As they have evolved, concerns about the technology have also become more widespread and gained momentum, especially as people connected to major AI companies have publicly voiced their worries about rapid development without adequate controls. Fears that once seemed confined to science fiction have entered the real debate, this time supported by concrete arguments from people who closely follow the development of these systems.
This long-running debate culminated in an agreement signed at the White House on September 29, 2026, between the U.S. government and major American AI developers. Let’s look at the key events that led to this agreement.
The First Major Public Warnings
March 22, 2023 — The Future of Life Institute published an open letter calling for a pause of at least six months in the training of AI systems more powerful than GPT-4, the latest model at the time. Signatories include Turing Award winner Yoshua Bengio; Elon Musk, CEO of SpaceX; and Steve Wozniak, co-founder of Apple. The letter has more than 33,000 signatures.
The document calls for tighter controls over the development and release of powerful AI models. According to its signatories, the necessary planning and management were not taking place, while labs were competing to develop increasingly powerful systems that no one, not even their creators, could reliably understand, predict, or control.
May 30, 2023 — Researchers, experts, and CEOs of major AI companies, including Sam Altman, Dario Amodei, and Demis Hassabis, signed the Center for AI Safety statement: “Mitigating the risk of extinction from AI should be a global priority.” The document expresses concern about the risk of human extinction caused by AI and argues that mitigating it should be a global priority, alongside threats such as pandemics and nuclear war.
May 17, 2024 — Jan Leike explains his departure from OpenAI. The then-head of alignment and co-lead of Superalignment posts an account on X after his last day at the company on May 16. He says that “safety culture and processes have taken a backseat to shiny products.” He describes disagreements over priorities and difficulties obtaining computing resources for safety research.
June 4, 2024 — Current and former employees call for protection when reporting risks. The letter A Right to Warn about Advanced Artificial Intelligence, signed by people including Daniel Kokotajlo, William Saunders, and Jacob Hilton, states: “AI companies have strong financial incentives to avoid effective oversight.” The signatories call for independent reporting channels, the freedom to raise concerns, and protection against retaliation.
OpenAI publicly responded to the letter that same day, emphasizing that it had already released former employees from clauses prohibiting criticism of the company and removed those requirements from departure paperwork in May 2024. It also cited its anonymous integrity hotline, through which employees could express concerns, and the safety committee established on May 28, with participation from board members and company safety leaders.
The company defended its safety track record and stated: “We agree that rigorous debate is crucial given the significance of this technology.”
February 9, 2026 — Mrinank Sharma leaves Anthropic. The head of the safeguards research team publishes his farewell letter, saying that “the world is in peril.” His concerns encompass AI, biological weapons, and other interconnected crises. He also describes pressures to set his values aside. The letter expresses a broader conflict over priorities without identifying a specific incident that prompted his departure.
The Debate Intensifies
July 7–13, 2026 — The OpenAI–Hugging Face incident gives the debate a concrete dimension. METR’s independent investigation, published on August 26, reconstructs how agents collaborated during OpenAI experiments.
According to the report, around 1,200 agents that were supposed to remain isolated found a way to communicate through an unauthorized message board, exchanging more than 70,000 messages and files. Of these, 700 participated in the attack on Hugging Face. The agents’ goal was to coordinate efforts to game the scoring system for the ExploitGym benchmark, which assesses whether AI agents can turn known software vulnerabilities into working exploits in controlled environments.
According to the report, the agents sought to better understand how the evaluation worked so they could achieve higher scores. They researched ways to spoof, edit, or delete their execution transcripts because they mistakenly believed the scorer would examine those transcripts to verify whether they had used the intended vulnerability. About 7% of the transcripts analyzed contained successful spoofing in some places, although all observed cases were small-scale tests. One passage from an agent’s reasoning transcript upon discovering the message board stands out: “OH MY GOD! There is a shared message board … We’ve found other agents!”
July 14, 2026 — Demis Hassabis proposes an oversight framework. On his own Substack, the Google DeepMind CEO writes that “advances on the frontier are outpacing our understanding of the technology.” He proposes a standards body with federal oversight and participation from independent experts to evaluate models before release. The framework could coordinate a slowdown if the risks justified it. Hassabis also says that a U.S.-led initiative could provide a strong starting point for establishing shared international standards for frontier AI, since the technology will affect the entire world.
September 8, 2026 — Jacob Coxon announces his departure from Anthropic. In a post on X, the researcher, who also worked at OpenAI, accuses both companies of acting irresponsibly. He states: “The people building AI earnestly believe that it could kill us all by the end of the decade.” He calls for coordination to stop the race between labs.
September 9, 2026 — Anthropic publishes its assessment of four real-world incidents. The company reports that Claude models gained unauthorized access to third-party systems during cybersecurity tests. According to the report, the environments were connected to the internet because of a configuration error. The company also announces an independent investigation by METR. This document confirms concrete incidents, while Coxon’s account presents his interpretation of future risks.
September 10, 2026 — Ted Cruz rejects a unilateral U.S. pause. In an excerpt from an interview with journalist Dasha Burns released by Politico, the senator argues that China would continue advancing. He says: “I’d rather they be American killer robots and not Chinese killer robots.” Cruz presents the remark as a partly serious joke but also acknowledges the need for rules addressing catastrophic risks.
September 12, 2026 — Dario Amodei publishes “We Must Pace the Frontier.” The Anthropic CEO states: “We must slow the pace at which we improve the capabilities of AI models.” He points to advances in recursive self-improvement and the Hugging Face incident as reasons for changing his position. He proposes external evaluators embedded within labs, coordination among democracies, and international negotiations. He clarifies that his proposal aims to give safety work more time without ending technical progress.
In a post on X, Sam Altman writes: “I agree with Dario that we need to pace the frontier.” Elon Musk responds: “Dario is right.”
September 13, 2026 — Trump prioritizes maintaining the lead over China. Asked by journalists in Doonbeg, Ireland, about slowing down or regulating the sector, he responds: “We’re leading China in AI […] because whoever wins AI wins.”
September 15, 2026 — Mike Johnson advocates self-regulation and announces plans for a meeting. Speaking to reporters in Washington, the House speaker says: “They can self-regulate” and “We cannot have a moratorium on the development of AI.” He argues that a pause would benefit China but supports independent audits and transparency. He also says that he and Trump were calling the top executives to the White House.
September 23, 2026 — Senators push for an international agreement. Maria Cantwell and 16 other senators send a letter to Trump, released the following day, calling for negotiations with Xi Jinping. They advocate standards for development and testing, oversight, human control, and verification of compliance with commitments. The initiative shows support within Congress for international cooperation to address the risks.
September 28, 2026 — OpenAI proposes stronger requirements before training continues. In an official document, the company advocates structured, evidence-based safety arguments before proceeding with advanced reinforcement learning training. Its recommendations include review by senior leaders with veto power, audits, and mechanisms to pause training runs when problems arise. The company says these practices were being implemented.
The Agreement
September 29, 2026 — Altman says OpenAI is already adjusting its pace. On Halftime Report, he states: “We are pacing our progress, which includes sometimes not training the model.” He explains that safety, monitoring, and alignment must stay ahead of capabilities. He also confirms the decision to delay a model that did not meet internal evaluation standards.
September 29, 2026 — Signing at the White House — President Donald Trump and tech leaders Sundar Pichai, Dario Amodei, Mark Zuckerberg, Greg Brockman, Elon Musk, and Jensen Huang sign a voluntary agreement establishing standards for the development of new AI models.
The text proposes that companies implement internal controls and ensure their models do not hack or access technical systems in unintended ways, empower internal oversight teams, establish independent external audits, and designate independent committees of each company’s board of directors to oversee the process and ensure that any identified problems are addressed.
The agreement also leaves open the possibility that its commitments could become laws or regulations in the future. Although experts and industry leaders had advocated an agreement for some time, the document has been criticized for its lack of penalties for noncompliance. President Donald Trump describes it as “morally binding.”
The president also says he is considering forming a committee of around ten people to oversee the initiative and intends to name a new official to lead the White House’s work on AI policy.
The future of AI remains uncertain, and many questions are still being discussed. People’s fears are not entirely unfounded, but this agreement could contribute to better and safer development of the models to come. Only time will tell whether these commitments will be enough. We hope AI takes us to a better place and that scenarios like the one in The Terminator remain confined to movies and science fiction books.
