At 3:12am your agent decides the fastest way to “clean up stale customer records” is to archive 1,400 rows in your CRM. It has the OAuth scope to do that. Nobody is awake to see it. By morning the action is done, the audit trail is whatever the vendor chose to log, and your first question is one no one can answer: who approved that?

That scenario isn’t hypothetical anymore. At DevDay 2026 in San Francisco on September 29, OpenAI launched “dots”: always-on agents that get their own cloud computers and can reach roughly 4,000 app connections. Next to GPT-6.1 Sol and the new enterprise Marketplace, dots was the part of the keynote that changes how teams need to think about AI agent security. A chatbot waits for a prompt. A dot keeps working toward a goal while you sleep.

We build AI features for a living, and we start with security instead of adding it at the end. Our view is simple. Persistent agents with standing access to thousands of apps turn every OAuth scope and every API key into attack surface. The model matters less than the question of what it’s allowed to touch, and who finds out when it does.

Why always-on changes the threat model

Most LLM integrations shipped in the last two years were request/response. A user asks, the model answers, and maybe it calls a tool or two inside a session that ends. The blast radius was limited by time and by the fact that a human was watching.

Always-on agents remove both limits. Three things change at once:

  • Credentials become standing, not session-bound. An agent that works overnight needs tokens that last overnight. Long-lived tokens with broad scopes are exactly what attackers go looking for.
  • Every input becomes a possible instruction. An agent reading your inbox, your tickets and your shared docs is reading text written by strangers. Prompt injection stops being a demo trick. It becomes a real way for an outside email to steer an agent that holds write access to your billing system.
  • Connections multiply quietly. Four thousand available integrations means the default path is “connect it, see if it helps.” Each connection is another scope grant, and those almost never get revoked.

None of this is a knock on OpenAI specifically. The same logic applies to anything built on the OpenClaw Enterprise control plane, which the OpenClaw Foundation (Red Hat, Nvidia and OpenAI) open-sourced the same week. An open control plane is good news, because you can inspect it, self-host it and put your own policy in front of it. But open code doesn’t secure itself. A control plane is only as strict as the policies someone actually writes into it.

An AI agent security baseline before anything runs overnight

Whether you adopt dots, build on OpenClaw Enterprise, or roll your own agent loop in Python or Node.js, we’d want these four controls in place before an agent gets unsupervised hours:

  1. Scoped, short-lived credentials per agent, per task. Don’t hand an agent your admin OAuth grant. Issue it a dedicated identity with the narrowest scopes the goal needs: read-only where possible, specific resources rather than whole workspaces. Tokens should expire and be re-issued, not live forever in a config file. If the agent’s job is summarising support tickets, it should be physically unable to issue refunds.
  2. An action log you own. Vendor dashboards are useful, but your audit trail shouldn’t depend on another company’s retention policy. Record every tool call on your side: the goal, the input that triggered it, the exact API request, and the response. When something goes wrong at 3am, you need a replayable record, not a summary.
  3. Human approval for anything irreversible. Sort actions into tiers. Reading is free. Drafting is free. Sending an external email, deleting data, moving money, changing permissions or deploying code waits in a queue until a person approves it. The agent can prepare the change overnight, and a human clicks “yes” at 9am. That costs you a few hours of latency. It saves you an incident review.
  4. Treat retrieved content as untrusted input. Anything the agent reads from an email, a web page or a shared doc should be clearly separated from its instructions, and tool calls triggered by that content should face stricter policy. This is ordinary application security thinking, the same instinct that makes you parameterise SQL, applied to a new kind of interpreter.

Put one more on the list: a kill switch that actually works. You should be able to revoke one agent’s credentials in seconds without breaking every other integration. If you can’t, your identity architecture is too coarse for agents.

The unglamorous work is the real work

Nobody puts “we wrote a permission matrix” on a keynote slide. Still, the gap between an agent pilot that impresses leadership and one that survives its first security review is almost always backend and architecture work: an identity layer, a policy check in front of every tool call, an append-only log, and an approval queue with a small admin dashboard so a human can see what’s pending and why.

That’s regular software engineering. It’s API development, database design and access control, the same things that decide whether any production system is trustworthy. Agents don’t make that work optional. They make it matter more, because the thing calling your APIs never gets tired and never second-guesses itself.

Our honest take on DevDay: dots and GPT-6.1 Sol will make agents capable enough that teams will want to give them real responsibility. That’s the right moment to slow down just enough to decide which responsibilities, under which credentials, with whose sign-off. Capability is arriving faster than most organisations’ access controls. The teams that do well won’t be the ones that connect the most apps. They’ll be the ones that can say exactly what their agents did last night, and prove it.

Where we come in

At orithLabs we build AI integrations with security as the first design constraint, and we audit existing codebases where AI features were bolted on quickly and now need proper guardrails. If you’re weighing dots, experimenting with OpenClaw Enterprise, or just trying to work out what your current agent can actually reach, we’re happy to talk it through engineer to engineer. No pitch deck, just a straight answer about where the risk sits.