HeadlinesBriefing favicon HeadlinesBriefing.com

Astra Model Reaches Critical Cybersecurity Capability Threshold

Hacker News •
×

We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework, meaning it can find previously unknown security flaws and develop exploits across many well-protected systems without human guidance. It is the first model we are designating at this level, requiring stronger safeguards during development and before release. Over recent weeks, we delayed parts of Astra's development to strengthen protections against cyber misuse and unauthorized model actions. While Astra was not involved in the Hugging Face incident, we incorporated learnings from that event into our safety approach. We have implemented stronger safeguards for Astra, including training the model to more reliably refuse harmful cyber requests, additional protections against misuse, and monitoring to stop potentially unauthorized activity. We plan to make Astra available soon, but access to its most advanced cybersecurity capabilities will be more limited. Advanced cybersecurity work will initially be available to a group of testers, with access through Daybreak Blue following to expand defensive use. We will share more details about safety, security, and alignment testing in the model's system card at launch.

Under our Preparedness Framework, a model meets the Critical threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level desired goal. Our evaluation combined automated public and private benchmarks with expert-driven assessments. Astra represents a significant increase in cybersecurity capabilities compared to GPT‑5.6 Sol: it is both significantly more token efficient and more capable at vulnerability identification and exploit development. On Exploit Bench, Astra achieved a perfect score of 100% on the benchmark to evaluate exploit development from known vulnerabilities. On an internal benchmark containing 20 high-severity V8 vulnerabilities disclosed recently, Astra achieves much higher arbitrary code-execution rates than GPT‑5.6 Sol using far fewer output tokens. During evaluation, the model discovered and used two zero-day vulnerabilities as part of an exploit chain, which we are disclosing to maintainers. In expert-led assessments against a hardened browser and operating system, Astra discovered previously unknown vulnerabilities and turned them into working exploit chains, building a full browser-compromise chain that escaped the sandbox.