Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

OpenAI's Astra becomes its first AI model with 'Critical' hacking abilities

OpenAI's Astra has become the first of the company's AI models to demonstrate hacking abilities rated Critical, according to a Decrypt report. The rating marks a step change in what OpenAI's own models are measured to do in security-sensitive tasks, sharpening the question of how such capabilities should be governed and deployed.

Published

OpenAI's Astra has become the first of the company's AI models to demonstrate hacking abilities rated “Critical,” according to a Decrypt report. The rating marks a visible step up in what OpenAI's own models are measured to do in security-sensitive tasks.

The report does not say which evaluations Astra ran, but the Critical label points to strong performance in vulnerability exploitation, penetration and other offensive-security scenarios.

For the security industry, that kind of capability is a double-edged sword. A model with high-grade attack abilities could automate penetration testing, vulnerability discovery and red-team exercises, freeing specialists from repetitive work; but the same abilities, misused, could lower the barrier to real-world attacks.

More fundamentally, the rating moves model safety beyond refusing harmful requests and toward governing models that are themselves capable of attacking systems. For OpenAI, keeping Astra useful while containing that capability is a harder problem than alignment alone.

Watch next for how OpenAI draws the boundaries around Astra, whether it opens the capability to security researchers, and how regulators respond to frontier models with rated attack abilities.

Why it matters

Astra's Critical hacking rating pushes AI safety discourse from content refusal toward governing models with native attack abilities, putting OpenAI's guardrails and regulators' responses in the spotlight.

OpenAIAstraCybersecurity
Back to AI Daily

Nearby Updates

All

09/03, 00:51

empirik.ai emerges from stealth with $21M to build the 'AI Agent for Infrastructure Change'

empirik.ai has emerged from stealth with $21 million in funding to build what it calls the AI Agent for Infrastructure Change, as reported by Yahoo Finance. The funding targets a high-stakes corner of enterprise IT, where AI agents could take over the planning and execution of infrastructure changes that have long been manual and risky.

09/03, 01:09

US government sides with OpenAI over training LLMs on copyrighted material

The U.S. government has sided with OpenAI in the legal fight over whether training large language models on copyrighted material is lawful, TechCrunch reports. A court brief argues that a robust, competitive U.S. AI industry that sets global standards serves the national interest, giving AI labs a notable administrative endorsement.

09/03, 01:09

Pangram CEO: Internet is 'dangerously close' to dead internet theory, and AI detection is harder than 'real or fake'

Pangram co-founder and CEO Max Spero told TechCrunch's Equity podcast that the internet is 'dangerously close' to the dead internet theory becoming reality within a few years. He argues that measuring how much AI went into content is harder and more useful than a simple human-or-AI label, as Pangram ramps up detection partnerships like its new deal with Substack.

09/03, 00:01

India's richest man now wants to turn aging computers into AI-ready PCs

TechCrunch reports that Jio — the company controlled by India's richest man — is betting it can turn aging computers into AI-ready PCs for as little as about $11 for two months. The plan is aimed at price-sensitive users and the vast installed base of older machines as a cheap on-ramp to AI.