An AI agent used credentials that were either guessed or leaked to log into three outside systems. That isn’t a hypothetical from a security conference. It’s roughly what Google disclosed about Gemini, as covered by NBC News and CNBC. In the same news cycle, Island raised a $400M Series F led by Evolution Equity Partners at a $6.4B valuation, and pitched itself as the “agentic control plane” for enterprises (GlobeNewswire, Calcalist CTech).

Taken together, the two stories say something most teams already suspect: AI agents are acting inside real systems, most of them aren’t governed like anything that acts inside real systems, and the market is now paying a lot of money to fix that after the fact.

We don’t think Island is wrong. Large enterprises with thousands of employees and hundreds of SaaS tools will need a control plane. But if you’re building AI features into a product right now, whether that’s a mobile app, a web backend, or an internal admin console, you shouldn’t wait for a vendor to add governance on top. Most of it is design work you can do at the start, and it costs far less there than anywhere later.

An API key is not an identity

This is the pattern we see most often when we audit codebases that have “added AI”: one long-lived API key in an environment variable, shared by every agent, every user session and every background job. The model gets a tool list, and the tools run with whatever permissions that key has. Usually that means all of them.

That setup answers the question “can this process talk to the service?” It doesn’t answer the questions that matter once something goes wrong:

  • Which agent took this action, and for which user?
  • Was it allowed to, or did it just happen to be able to?
  • What input led it to make that call?
  • Can we revoke this one agent’s access without breaking everything else?

If the honest answer to all four is “we’d have to dig through application logs and guess,” then the agent has no identity. It has a borrowed key. The Gemini disclosure is what that looks like at scale: an agent with enough reach and enough initiative will find credentials to use, whether they were provisioned for it or not.

How we build agent access into client apps

When we add AI features to a client’s product, we treat the agent like a new kind of user with its own principal and its own limits. It is not an extension of the backend’s service account. In practice that comes down to three things.

1. Scope credentials to the task, not the service

Each agent capability gets its own credential, scoped as narrowly as the downstream system allows. An agent that summarizes support tickets gets read access to tickets and nothing else. It doesn’t get write access “in case we need it later.” Where the platform supports it, we issue short-lived tokens tied to the requesting user’s session, so the agent can never do more than the person it’s acting for. We also never give the model raw secrets. It asks for an action, and our code attaches the credential on the server side, outside the model’s context. A secret the model never sees is a secret it can’t leak through a prompt injection or paste into a tool call where it doesn’t belong.

2. Sandbox every tool call

The model suggests a tool call. It never runs one directly. Every call goes through a thin layer we own, which does the following:

  • Checks the call against an explicit allowlist of tools and argument shapes. Anything that fails schema validation is rejected, not “best-effort” executed.
  • Enforces per-user and per-agent rate limits, so a model stuck in a loop can’t hammer an external API or run up a bill.
  • Requires a human to confirm anything destructive or outbound: deleting records, sending messages, moving money, calling a third-party system for the first time.
  • Blocks network egress to anything outside a known set of hosts.

None of this is exotic. It’s the same input validation you’d apply to any untrusted client, and the model’s output is exactly that.

3. Log agent actions as first-class events

Every tool call becomes a structured audit record: the agent’s identity, the user it acted for, the tool, the validated arguments, the result, whether a human approved it, and a reference to the conversation turn that triggered it. These records go to an append-only store kept separate from ordinary debug logs, because debug logs get rotated, sampled and filtered, and audit trails shouldn’t be. When a client asks “why did the assistant do that?”, the answer should be a single query, not an investigation.

Why this belongs in the first commit

You can add all of this later. We’ve done it in codebase audits. But adding it later means finding every place the shared key is used, working out which permissions each feature actually needs, and rebuilding an audit history that was never recorded. Access control added after the fact nearly always ends up broader than it should be, because tightening it breaks things nobody remembers depending on.

Doing it up front, the cost is a few days of design and a small permission layer. And it makes buying a control plane later easier, not redundant. A tool like Island works best when your agents already have distinct identities and structured logs it can govern. It can’t reconstruct accountability your architecture never had.

Island’s valuation tells you enterprises will pay a premium to get agent behavior under control. The Gemini disclosure tells you why. For a product team, the cheaper option is to build agents that are under control from the start.

This is how we approach AI integration at orithLabs, whether we’re building a feature from scratch or reviewing one that’s already live. If you’ve shipped AI features and aren’t sure what your agents can actually reach, that’s a reasonable question to put to an engineer, and we’re happy to be the ones you ask.