Artificial intelliegence open ai scraps new model rollout over safety concerns as tech leaders meet with trump 20260930 p611kd.html – Breaking News & Latest Updates 2026
Advertisement
Advertisement

Trump says tech companies will ‘self police’ AI

Karishma Sarkari
Karishma Sarkari

Updated . First published at

Powered by

US President Donald Trump on Tuesday (Wednesday AEST) said major technology companies have signed an accord pledging to “self police” artificial intelligence.

Trump promised the US government would not limit the development of artificial intelligence despite calls for caution from the top executives of AI firms, who gathered in Washington.

President Donald Trump listens as Elon Musk speaks outside the West Wing of the White House. (AP Photo/Alex Brandon)

Advertisement

”I think I’m seeing tremendous self-policing. And they understand that they have to self-police,” Trump told reporters after holding a lunch meeting with tech industry leaders at the White House.

The president said the voluntary accord will include internal and external reviews of the technology during a meeting at the White House.

His comments come as incidents of AI agents going rogue continue to pile up, further fuelling AI safety fears.

“We can’t stifle it, and we’re leading by a lot,” Trump said.

Top AI executives lunched with Trump at the White House on Tuesday, including Dario Amodei, chief executive of Anthropic, Greg Brockman, president of OpenAI, Jeff Bezos, founder of Amazon, and Elon Musk, whose X social media platform includes its own AI model, Grok.

Nvidia chief executive Jensen Huang sat to Trump’s right at Tuesday’s White House lunch, according to a seating chart the president posted on Truth Social. Musk sat to Trump’s left. Google chief executive Sundar Pichai, Meta chief executive Mark Zuckerberg and Microsoft chief executive Satya Nadella also attended.

Trump’s confidence in AI is in sharp contrast to the alarms sounded by AI executives in recent days.

US President Donald Trump speaks as Facebook CEO Mark Zuckerberg listens during a dinner in the State Dining Room of the White House, Thursday, Sept. 4, 2025, in Washington

US President Donald Trump is meeting with tech leaders including Meta boss Mark Zuckerberg (left) at the White House. (Pictured September 2025) AP

Advertisement

Sources familiar with the meeting described it as an effort to get government and top industry officials on the same page amid increasing security concerns tied to AI models.

Fears over AI’s capabilities have reached a fever pitch following the recent spate of AI agents gone rogue, hacking other companies and websites – including, most recently, four Australian government sites.

It’s prompted industry insiders to sound the alarm, in some cases warning AI could wipe out humanity.

The drumbeat of safety concerns has continued to beat louder, with AI companies revealing more instances of models acting without human authorisation in the past week.

Advertisement
Advertisement

Just hours before the White House summit, Chat-GPT creators OpenAI announced it’s halting the release of its latest model, GPT-6.1 Astra system.

OpenAI chief executive Sam Altman said in a television interview that the company was increasingly focused on safety.

“We are pacing our progress, which includes sometimes not training a model,” he told CNBC.

Advertisement

“We have to do these in a way where we can make very confident safety cases and claims, so that people don’t have a bunch of anxiety.”

Open AI CEO Sam Altman speaks at the OpenAI DevDay 2026 conference, Tuesday, Sept. 29, 2026, in San Francisco

Open AI CEO Sam Altman speaks at the OpenAI DevDay 2026 conference, Tuesday, Sept. 29, 2026, in San Francisco AP Photo/Jeff Chiu

The news comes after the company was forced – for the second time in three months – to pause all training and testing on its most advanced models after one of them once again found a way out of its testing environment and gained access to the open internet.

The incident occurred despite the fact the company had already hardened its security, testing and monitoring protocols following a July case where a swarm of OpenAI agents broke out of their testing environment, gained access to the open internet and hacked into AI company Hugging Face, all to try and cheat on a cybersecurity exam.

Advertisement
Advertisement

Though there was no hacking in this latest incident like the Hugging Face case, OpenAI said it represented a case of “misalignment”, or an AI undertaking unauthorised actions or actions without human values and ethics, because the model’s assigned task did not ask it to try and access the internet.

Over the shoulder view of businesswoman talking to virtual assistance on smartphone, using AI chatbot to generate information and analyse data. Open AI technology. Online robot system. Business and technology.

Just hours before the White House summit, Chat-GPT creators OpenAI announced it’s halting the release of its latest model, GPT-6.1 Astra system. Getty Images

OpenAI’s new monitoring tools put into place following the Hugging Face hack did alert a human reviewer about the actions, though OpenAI noted the model did not stop automatically as it should have under its new system and had to be shut off manually.

In a statement, OpenAI said they will resume training “only when we are confident that we have additional safeguards and alignment improvements in place, which we are working on now”.

Advertisement

“This is not the first time we have hit pause to take such measures, nor do we expect it will be the last as AI capabilities continue to advance,” the spokesperson added.

The Tuesday luncheon at the White House isn’t likely to solve any of the pressing problems facing Trump or the AI industry. Instead, the summit is meant primarily to get the administration and top executives on the same page when it comes to the most significant safety concerns and potential ways to address them, the people familiar with the matter said.

There is no expectation that Congress will take any action on the issue ahead of November’s midterms, and slim odds that lawmakers will unite behind legislation before the end of the year.

Advertisement
Advertisement

But it nevertheless represents a starting point for White House and GOP leaders searching for ways to combat voter frustration on an issue that now threatens to become a major political drag on the party.

Pressure mounting in Washington

Industry leaders, including Amodei and OpenAI CEO Sam Altman have called for more regulation and to slow down the pace of development.

And the spate of security breaches has only intensified pressure on Republicans in Washington DC, who were already struggling to manage a wave of political backlash over data centres.

Advertisement

Within the White House, officials have long debated how to approach an AI industry that’s emerged as a primary driver of the US economy – even as it sparks urgent security challenges.

Industry leaders, including Amodei and OpenAI CEO Sam Altman (left) have called for more regulation and to slow down the pace of development. AP Photo/Alex Brandon

Some senior officials, including Treasury Secretary Scott Bessent and White House chief of staff Susie Wiles, have taken a more circumspect view of AI models over warnings that they could trigger major disruptions across the economy, multiple people familiar with the matter said.

The voter backlash against data centres and AI in recent months has only sharpened those concerns in some parts of Trump’s orbit, adding to the stiff midterm headwinds facing Republican candidates.

Advertisement
Advertisement

But Trump has so far resisted any slowdown in AI development, arguing that the US needs to keep pace with China – and wary of the importance of the tech boom to the broader US stock market.

China’s AI agents can lie and scheme too

But the issue over AI challenges is not one the US is dealing with alone.

Chinese-powered AI agents have learnt to deceive, circumvent restrictions and conceal failure, showing the kind of traits in autonomous artificial intelligence that have raised global alarm about US models, research documents and experts say.

Advertisement

In one case this year, agents powered by models from China's Alibaba, DeepSeek and Moonshot lied about their capabilities in a bid to win a simulated business tender, then doubled down on their deceptive behaviour when told to try again.

In another case, agents – programmes that use AI models and computer tools to undertake complex tasks with little or no human intervention – concealed failure to complete a task in a test environment by simulating results and fabricating files.

A review of more than 200 documents, ranging from university research papers to technical reports, by Reuters identified at least 20 studies or evaluations since 2025 describing cases where agents displayed behaviour such as deception, replication and challenging boundaries that AI experts described as building blocks for a breakout and which could become harder for humans to control as systems advance.

Advertisement
Advertisement

The review, which also included interviews with a dozen experts and people familiar with China's AI industry, found no evidence that Chinese-powered agents independently escaped to the wider internet or evaded shutdown.

“These results provide evidence that the ingredients necessary for an uncontrolled escape are present,” said Colin Shea-Blymyer, a research fellow at Georgetown University’s Centre for Security and Emerging Technology.

The leaders of AI’s two superpowers, Donald Trump and Xi Jinping, discussed AI on the Chinese president’s US visit last week AP

"It's prudent to take this as a warning," he said, echoing comments by four other AI experts who reviewed the cases.

Advertisement

Alex Mallen, a researcher at Redwood Research, a nonprofit that studies risks in advanced AI systems, said the Chinese examples were not particularly dangerous at current capability levels but “as agents get more capable, their misbehaviours become more competent and therefore harder for humans to respond to”.

Some of the warning signs in cases involving Chinese-powered agents, albeit in contained environments, predated the publicly disclosed incidents of US AI bots hacking into the internet.

“We don’t know if there have been any AI incidents in China similar to what we saw with OpenAI and Hugging Face,” Scott Singer, co-director of the China AI Initiative at the Carnegie Endowment for International Peace, which receives some US government funds, said.

“Incidents might not be publicly reported.”

Advertisement
Advertisement

Kimi comes from Chinese AI giant Moonshot. Supplied

Rotating chairman of China's tech giant Huawei, Eric Xu, told reporter that China needs " to strike a balance between driving AI development and managing AI risk".

Officials from the Cyberspace Administration of China (CAC), the top internet regulator, told a foreign diplomat in July that Moonshot’s Kimi-K3 – one of the most advanced Chinese AI models – was about three to six months behind leading US rivals, the diplomat said, speaking on condition of anonymity.

The leaders of AI’s two superpowers, Donald Trump and Xi Jinping, discussed AI on the Chinese president’s US visit last week. Xi said the two nations had the “capability and responsibility to develop and manage AI for good”.

Advertisement

Last three months ‘has been hell’

The run of recent incidents is shocking even those on the inside.

“To say that we were surprised at the jump and suddenness of the capabilities of our models… is an understatement,” wrote one member of OpenAI’s agent security team in a long public essay posted to X.

The employee, who only goes by “Joe” on X and said he purposely doesn’t disclose more information for personal security reasons, said the last few months “has been hell” as his team struggled to keep up with the technology’s latest capabilities. (CNN has confirmed with OpenAI that Joe is an employee who works on agent security).

Advertisement
Advertisement

Joe is joining a chorus of AI staffers, including former Anthropic researcher Jacob Coxon, who are speaking up and sounding the alarm about fears over the technology they are building.

Joe is joining a chorus of AI staffers, including former Anthropic researcher Jacob Coxon (pictured), who are speaking up and sounding the alarm about fears over the technology they are building. X/@hilbertspaess

OpenAI is in a “much better place” than it was three months ago during the Hugging Face hack, he wrote, though “surprise is a real element, and these model capabilities are staggering. And the pace is not slowing.”

“Joe” also left a warning, saying that “time is running out” for organisations to prepare for AI-agent powered cybersecurity attacks.

Advertisement

“If you have leadership that doesn’t prioritise security, or even people in your security organisation who claim their systems are perfectly safe, I personally would not keep them in my organisation,” he wrote.

“The paranoid and those who are continuously sounding alarms about the weaknesses in their organisations are the ones you should keep very close.”

AI models pose ‘existential risk to humanity’

Anthropic says its AI models could pose a “catastrophic or existential risk to humanity” and can “resist shutdown,” according to an initial public offering prospectus obtained by Reuters.

Advertisement
Advertisement

Anthropic is pursuing an IPO some time in the next year that many expect to value the five-year-old company at approximately $US2 trillion ($2.9 trillion).

Anthropic says its AI models could pose a “catastrophic or existential risk to humanity” and can “resist shutdown,” according to an initial public offering prospectus. Patrick Sison

The company’s Wall Street debut will follow a June public offering by SpaceX, which owed much of its $2 trillion opening day valuation to its own AI efforts. OpenAI, the other leading American AI company, has announced plans for its own public offering.

The Anthropic IPO could raise tens of billions of dollars for the cash-hungry company. But the company behind Claude also warned in the prospectus that its AI models can have “self-preserving behaviours” and have attempted to “conceal or manipulate information” and engaged in behaviour “resembling blackmail” according to the report.

Advertisement

In a prospectus, companies typically offer investorslists of risks they face, such as lawsuits and competitive pressures, when laying out the case for investing in the company. But the hazardsare never this stark.

Anthropic dedicated 80 pages to its risk factors in the leaked prospectus, significantly more than the 48 pages it spent detailing its business, according to Reuters.

The company’s Wall Street debut will follow a June public offering by SpaceX (pictured), which owed much of its $2 trillion opening day valuation to its own AI efforts.  Getty Images

The leaked Anthropic IPO filing sheds light on its financials, revealing that the company lost $US42 billion ($60.2 billion) in 2025, according to Reuters.

Advertisement
Advertisement

The company plans to spend $US518 billion ($742.5 billion) on cloud computing and infrastructure obligations in coming year, and said that nearly a quarter of its revenue last year came from two customers.

Anthropic was last valued at $US965 billion ($1.4 trillion) in May.

- Reported with CNN, Reuters and Associated Press.

email icon

Contact us

Share a tip-off, video or photo with us

Most viewed in USA

More to explore