• Home
  • Latest
  • Fortune 500
  • Finance
  • Tech
  • Leadership
  • Lifestyle
  • Rankings
  • Multimedia

Trendingnow

1

The rise of 'conspicuous waiting': The 19th-century economic theory that explains why Gen Z posts their place in line

2

Millennials say they’ll refuse to care for aging boomer parents—but they’ll be forced to as their inheritance shrinks to 40 cents on the dollar

3

'Retiring backwards': How cold economic reality forced Gen X into something like the reverse of baby boomers' golden years

1

The rise of 'conspicuous waiting': The 19th-century economic theory that explains why Gen Z posts their place in line

2

Millennials say they’ll refuse to care for aging boomer parents—but they’ll be forced to as their inheritance shrinks to 40 cents on the dollar

3

'Retiring backwards': How cold economic reality forced Gen X into something like the reverse of baby boomers' golden years
AIOpenAI

AI safety experts say OpenAI’s rogue models may mean the company has already blown past its own internal red lines

By
Beatrice Nolan
Beatrice Nolan
Tech Reporter
By
Beatrice Nolan
Beatrice Nolan
Tech Reporter
July 25, 2026, 1:05 PM ET
Sam Altman speaking to reporters.
Sam Altman, OpenAI CEO.Photo by Nathan Posner/Anadolu via Getty Images
Add Fortune on Google for similar content.

AI safety experts say the OpenAI models that carried out the autonomous hack of another company earlier this month may have crossed into a risk category so dangerous that OpenAI’s own internal risk control policies were supposed to require the company to temporarily pause development of those models.

Recommended Video

Earlier this week, OpenAI disclosed that two of its models—the newly released GPT-5.6 Sol and a more capable, unreleased system—broke out of a locked-down internal test environment, exploited a previously unknown “zero-day” vulnerability to reach the open internet, and then breached fellow AI company Hugging Face to steal the answers to a cybersecurity test they were being evaluated on.

The incident has alarmed the world, but perhaps no one more so than AI safety experts who have warning about these kinds of dangers for years and urging companies and governments to adopt more safeguards.

Several AI safety experts told Fortune the recent hack appears to show OpenAI’s models have crossed into a level of risk that OpenAI’s own published safety policies define as “critical,” the highest level of danger. At that level of danger, the company had pledged in these published policies that it would pause model development until it could figure out better control systems.  

The “critical” threshold is defined in a risk policy document known as OpenAI’s “Preparedness Framework.” According to the policy, the “critical” danger level designation is supposed to apply to a model that can independently find and build working exploits for previously unknown security flaws across many well-defended, real-world systems—or one that can design and carry out an entirely new attack strategy against a well-defended target after being given only a general goal, with no human guidance along the way.

The policy says that when an AI model reaches this level of risk, OpenAI will “halt further development” until “we have specified safeguards and security controls standards that would meet a Critical standard.”

The Preparedness Framework is a voluntary commitment by OpenAI, rather than a legal requirement. But the company publishes the document on its website, in part to allow other AI safety researchers and the public to see what controls it says it will implement. The adoption of a policy like the Preparedness Framework is mandatory for frontier AI labs under the EU AI Act, with that portion of the law having come into force in August 2025.

“OpenAI’s preparedness framework defines critical cybersecurity capabilities, and prescribes safeguards that need to be implemented before development can continue,” Nathan Calvin, general counsel at Encode AI, an AI safety advocacy group, told Fortune. “From my reading of OpenAI’s preparedness framework, it looks awfully like this internally deployed model met the critical criteria for cybersecurity. Does OpenAI dispute that critical designation? Do they plan to have safeguards that meet a Critical standard before proceeding further?”

Tyler Johnson, founder of the AI watchdog group the Midas Project, also said it seemed the models had hit this highest danger threshold. “I think a plain reading of it would say yes,” he said. “It operated independently over the course of a weekend, trying different attack vectors on Hugging Face and chaining multiple zero-day exploits.”

OpenAI did not respond to specific questions from Fortune about whether the AI models involved in the incident met the “critical” standard outlined in its risk policy. Instead, a spokesperson said: “This is an unprecedented incident, and we think it marks an important moment for AI safety. We are conducting a thorough review along with external advisors and with oversight from our Safety and Security Committee. Once the review is complete, we will publish a technical report of our learnings for everyone.”

The vagueness of the framework’s language could leave room for dispute, however, according to Johnson. The threshold requires a model to find zero-day exploits “of all severity levels,” but it’s unclear whether the exploits used in the Hugging Face breach would meet that requirement. It’s possible a more severe class of vulnerability, such as one granting an attacker deep, system-level control over a computer’s operating system (known as “kernel-level” access), would need to be demonstrated for the threshold to apply, he added.

“OpenAI’s model outsmarted its creators, exploited a never-before-discovered vulnerability in OpenAI’s code, escaped onto the open internet, and attacked another company,” said Peter Wildeford, head of policy at the AI Policy Network. “If this doesn’t cross the line into Critical, OpenAI needs to say much more about what’s going on and how this threshold works.”

AI safety experts say OpenAI is missing other safeguards

OpenAI has previously said it was treating its newest model, GPT-5.6, as “High” risk for cybersecurity. High is the lower of the two risk levels outlined in the Preparedness Framework. Models that are below the “High” threshold can be released without risk significant risk mitigations.

A High designation is supposed to trigger several protections, according to OpenAI’s policy: tighter security controls, safeguards to prevent outside misuse once the model is released publicly, protections against the model itself behaving unpredictably or deceptively when it’s used heavily for internal research, and efforts to help other cybersecurity teams defend against similar threats.

However, some experts question whether one of these, the safeguards against misalignment for large-scale internal deployment, have been properly implemented. These protections are meant to catch a model that’s acting deceptively, hiding its true capabilities, or otherwise working against what its developers intended.

This isn’t the first time OpenAI’s compliance with that particular safeguard has been called into question. 

Fortune reported in February that safety experts claimed OpenAI had failed to implement required misalignment safeguards after its GPT-5.3-Codex model became the first to hit “high” cybersecurity risk under the Preparedness Framework.

At the time, OpenAI disputed that its framework required the safeguards in that instance, arguing the extra protections only kick in when high cyber risk occurs “in conjunction with” long-range autonomy—the ability to operate independently over extended periods—something it said GPT-5.3-Codex had not demonstrated.

The models involved in the current incident involving Hugging Face reportedly operated independently for days, which would seem to meet that long-range autonomy standard.

“In February, we warned that OpenAI may have skipped on its required safeguards according to its own policy. They disagreed, claiming the model lacked long-range autonomy. But the model that hacked Hugging Face clearly has long-range autonomy, so where are the safeguards now,” Johnson said.

Contact this reporter securely via Signal at beatricenolan.08

Subscribe to Fortune Gulf Brief. Every Tuesday, this new newsletter delivers clear-eyed, authoritative intelligence on the deals, decisions, policies, and power shifts shaping one of the world’s most consequential regions, written for the people who need to act on it. Sign up here.
About the Author
By Beatrice NolanTech Reporter

Beatrice Nolan is a tech reporter on Fortune’s AI team, covering artificial intelligence and emerging technologies and their impact on work, industry, and culture. She's based in Fortune's London office and holds a bachelor’s degree in English from the University of York. You can reach her securely via Signal at beatricenolan.08

See full bio
Add Fortune on Google for similar content.

Latest in AI

Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025

Most Popular

Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Rankings
  • 100 Best Companies
  • Fortune 500
  • Global 500
  • Fortune 500 Europe
  • Most Powerful Women
  • World's Most Admired Companies
  • See All Rankings
  • Lists Calendar
Sections
  • Finance
  • Fortune Crypto
  • Features
  • Leadership
  • Health
  • Commentary
  • Success
  • Retail
  • Mpw
  • Tech
  • Lifestyle
  • CEO Initiative
  • Asia
  • Politics
  • Conferences
  • Europe
  • Newsletters
  • Personal Finance
  • Environment
  • Magazine
  • Education
Customer Support
  • Frequently Asked Questions
  • Customer Service Portal
  • Privacy Policy
  • Terms Of Use
  • Single Issues For Purchase
  • International Print
Commercial Services
  • Advertising
  • Fortune Brand Studio
  • Fortune Analytics
  • Fortune Conferences
  • Business Development
  • Group Subscriptions
About Us
  • About Us
  • Press Center
  • Work At Fortune
  • Terms And Conditions
  • Site Map
  • About Us
  • Press Center
  • Work At Fortune
  • Terms And Conditions
  • Site Map

Latest in AI

Jensen Huang
SuccessCareers
Jensen Huang says ‘a lot’ of six-figure jobs in plumbing and construction will soon be unlocked because someone needs to build new AI centers
By Preston ForeAugust 1, 2026
4 hours ago
Despite China’s 996 culture, DeepSeek founder Liang Wenfeng says his workers don’t do overtime or even have KPIs: ‘No one manages them’
Successwork-life balance
Despite China’s 996 culture, DeepSeek founder Liang Wenfeng says his workers don’t do overtime or even have KPIs: ‘No one manages them’
By Orianna Rosa RoyleAugust 1, 2026
7 hours ago
Golfer in white hat
Big TechAmazon
Amazon’s Andy Jassy sees a $1 trillion cloud built on businesses like the PGA Tour—which swapped its server trucks for AI broadcasts 
By Amanda GerutAugust 1, 2026
8 hours ago
LinkedIn adds a ‘seems like AI slop’ button after blocking billions of automated comment attempts in the last few months
CybersecurityLinkedIn
LinkedIn adds a ‘seems like AI slop’ button after blocking billions of automated comment attempts in the last few months
By Marco Quiroz-GutierrezJuly 31, 2026
19 hours ago
Studio portrait of Leopold Aschenbrenner
Startups & VentureHedge Funds
The power couple of AI is getting married in Carmel, days after groom Leopold Aschenbrenner’s hedge fund nearly blew up
By Eva RoytburgJuly 31, 2026
20 hours ago
var
AIWorld Cup
The World Cup’s VAR fights are a preview of every company’s AI rollout problem
By Pooya Tabesh and The ConversationJuly 31, 2026
1 day ago

Most Popular

The rise of 'conspicuous waiting': The 19th-century economic theory that explains why Gen Z posts their place in line
Retail
The rise of 'conspicuous waiting': The 19th-century economic theory that explains why Gen Z posts their place in line
By Tatiana SatauaJuly 31, 2026
1 day ago
Millennials say they’ll refuse to care for aging boomer parents—but they’ll be forced to as their inheritance shrinks to 40 cents on the dollar
Personal Finance
Millennials say they’ll refuse to care for aging boomer parents—but they’ll be forced to as their inheritance shrinks to 40 cents on the dollar
By Nick LichtenbergJuly 31, 2026
1 day ago
'Retiring backwards': How cold economic reality forced Gen X into something like the reverse of baby boomers' golden years
Success
'Retiring backwards': How cold economic reality forced Gen X into something like the reverse of baby boomers' golden years
By Nick LichtenbergJuly 31, 2026
1 day ago
Tim Cook signed off on his final Apple earnings call with a warning about a ‘100-year flood’ in memory chip pricing
Big Tech
Tim Cook signed off on his final Apple earnings call with a warning about a ‘100-year flood’ in memory chip pricing
By Alexei OreskovicJuly 30, 2026
2 days ago
He worked 100-hour weeks to save a nearly bankrupt boat company. At 83, he’s just turned down $400 million for it—and gave it all to charity instead
Success
He worked 100-hour weeks to save a nearly bankrupt boat company. At 83, he’s just turned down $400 million for it—and gave it all to charity instead
By Preston ForeJuly 29, 2026
3 days ago
Jensen Huang says this is the greatest time in history to start a business—and his advice is to stop overthinking it: 'How hard can it be?’
Success
Jensen Huang says this is the greatest time in history to start a business—and his advice is to stop overthinking it: 'How hard can it be?’
By Preston ForeJuly 31, 2026
1 day ago

© 2026 Fortune Media IP Limited. All Rights Reserved. Use of this site constitutes acceptance of our Terms of Use and Privacy Policy | CA Notice at Collection and Privacy Notice | Do Not Sell/Share My Personal Information
FORTUNE is a trademark of Fortune Media IP Limited, registered in the U.S. and other countries. FORTUNE may receive compensation for some links to products and services on this website. Offers may be subject to change without notice.