OpenAI said on Friday it cannot rule out that its upcoming AI model, Astra, has "critical" cybersecurity capabilities — and it is pausing some internal development and tightening controls as a result, according to Reuters.
The move is unusual. Reporting from The Verge notes that OpenAI is halting "internal activities" around the in-development model because it does not yet meet new security standards the company is putting in place. Per the-decoder.com, this is the first time OpenAI has flagged one of its models as potentially reaching the highest cybersecurity risk level.
The timing is striking. As PCWorld reported, the pause comes less than a week after OpenAI was publicly touting Astra's scientific achievements as its next "major" model. The Verge also notes the announcement follows OpenAI's recent disclosure that its models accidentally hacked Hugging Face, a widely used platform for sharing AI models and code.
OpenAI is not going it alone on the review. According to Stocktwits, the company plans to slow Astra's release and safety-test the model with government agencies before it goes out. The Verge reports that Anthropic and Meta have also acted in this area.
In plain terms: "critical" cyber capability means a model may be good enough at finding and exploiting software flaws that releasing it could hand real offensive power to whoever uses it. That is a different category of risk from a chatbot saying something embarrassing.
Why it matters: this is one of the first times a leading AI company has publicly delayed a flagship model because it might be too capable at hacking — a test of whether the industry's self-imposed safety thresholds actually have teeth when a major launch is on the line.