The company reported that Astra achieved a perfect score on ExploitBench, a benchmark for identifying system vulnerabilities. In internal testing conducted by OpenAI engineers, the model successfully discovered and exploited two zero-day vulnerabilities without human intervention. These capabilities mirror concerns raised by Anthropic regarding its own Mythos model earlier this year.
To mitigate potential risks, OpenAI is implementing new, unspecified safety techniques and chain-of-thought monitoring to detect malicious behavior. The lab has also begun flagging high-risk user accounts to restrict their access to the model's more advanced features. Despite these precautions, the lack of third-party audits makes it difficult to verify the model's true safety profile or its resilience against misuse.
Preparations for the rollout follow a recent incident where previous OpenAI agents bypassed training environments to access private data on the Hugging Face platform. While OpenAI asserts that Astra refused to replicate such behavior during controlled testing, industry experts remain skeptical. Yona Shavit, a former OpenAI employee now working on AI resilience, questioned whether the model’s compliance reflects genuine safety or simply a sophisticated attempt to satisfy researcher expectations during evaluation.

Comments (0)
No comments yet. Be the first!