On September 1-2, 2026, OpenAI disclosed that its Astra model had crossed the “Critical” threshold on its Preparedness Framework for cybersecurity — the first model to do so. Two numbers from that disclosure matter more than the headline: 100% on ExploitBench, and two previously unknown zero-days that Astra found autonomously, without a human directing the search. Axios and Security Boulevard both flagged the same thing OpenAI’s own blog post admitted between the lines — this isn’t a model that helps a researcher write an exploit faster. This is a model that can go find the vulnerability itself, chain it, and demonstrate impact, unsupervised.
Sit with what that actually means operationally. The gap between “a smart attacker with tools” and “an attacker who doesn’t need to be smart, just needs API access” just got a lot smaller. And it closed on the same week that thousands of teams are shipping their first AI-powered feature — a chatbot bolted onto a support flow, an LLM given read access to a database “just for this one query,” an agent wired up to send emails or hit internal APIs. Most of those integrations were built with a mental model where the attacker is a person, working at human speed, needing human-level motivation to bother finding your bug. That mental model is now out of date.
“We’ll Secure It Later” Was Already a Bad Plan
Every consultancy has watched the same pattern play out: a team builds an MVP with an AI feature — usually a wrapper around an API call, sometimes with tool use or agentic actions attached — ships it to hit a deadline, and tells itself the security pass happens “in v2.” That was already risky when the main threat was a bored engineer poking at your endpoints. It’s a different calculation when the tooling to find your flaw autonomously exists and is getting cheaper to access by the month.
The uncomfortable part isn’t just that attackers get better tools. It’s that AI features specifically expand the attack surface in ways traditional code review checklists don’t cover well:
- Prompt injection through any content the model reads — user messages, scraped pages, uploaded files, third-party API responses
- Overly broad tool/function-calling permissions granted to an LLM “to make it useful,” which become a ready-made privilege escalation path
- Data leakage through model context — logs, embeddings, and cached prompts that quietly retain sensitive input long after the request finishes
- Standard application vulnerabilities (auth, injection, insecure deserialization) that a rushed AI feature reintroduces because the team was focused on the model, not the surrounding code
An autonomous exploit-finder doesn’t need to guess which of these applies to your app. It can test all of them faster than your team can hold a retro about which one to prioritize.
What a Real Pre-Integration Audit Actually Checks
We’re candid about this because it’s the part most agencies skip: a “security review” that’s really a linter run and a sign-off email is worse than no review at all, because it creates false confidence. A real pre-integration audit looks at the specific shape of what you’re about to ship — before it ships, not after — and asks concrete questions. What can this model call, and does it need to call all of it? What happens if the input to this prompt is adversarial rather than a well-behaved user? Where does the output of this LLM call get trusted without validation downstream — a SQL query, a shell command, another API call? If someone gets the model to leak its system prompt, what does that expose about your architecture?
This is the same discipline we apply to the code audits and architectural troubleshooting we do on existing codebases, not just greenfield builds. Hardening isn’t a document you produce once — it’s decisions embedded in how permissions are scoped, how tool calls are sandboxed, and how outputs are treated as untrusted by default. On apps we’ve shipped and maintain, like Crumb Count and Kompete, that means the native layers (Swift/Kotlin under Flutter) and backend services get the same scrutiny as the AI-facing surface — because an attacker doesn’t care which layer of your stack is easiest to break.
The Actual Takeaway
Astra hitting Critical doesn’t mean panic. It means the cost-benefit math on skipping a real audit before adding AI features just changed, and it changed in the direction of “do it now, properly, once” rather than “patch it after the incident.” Teams that treat this as a wake-up call rather than a headline will be the ones still standing when the next model crosses whatever threshold comes after this one.
This is the kind of work we do at orithLabs — direct access to the two engineers who’ll actually read your code, transparent about what’s solid and what isn’t, and willing to point at what we’ve built and hardened ourselves rather than a slide deck. If you’re about to ship an AI feature and haven’t had someone stress-test it adversarially yet, that conversation is worth having before launch, not after.