If your model can call refund_order(order_id, amount), it can call it with any order_id that ends up in its context window. That includes an ID someone pasted into a support ticket, a product review, or a PDF your retrieval step pulled in. The model doesn’t need to be jailbroken for this to happen. It only needs to be persuaded, and being persuaded is what language models do. LLM tool call security doesn’t come from a better system prompt. It’s an input validation problem, and the fix is the same one we use for any form submission from a browser we don’t control.

This post walks through the layer we put between a model’s tool-call JSON and the backend it touches, and explains why defending against prompt injection only at the prompt level fails when the executor trusts whatever the model emits.

The model is a client, not a component

The mistake we see most often when reviewing AI integration code is architectural. The LLM gets treated as part of the backend, a trusted internal service that happens to speak JSON. So its output gets routed straight into handlers that were written assuming the caller had already been authenticated, authorized, and validated.

The better mental model: the model is a client, like a browser or a mobile app. It sits on the far side of a trust boundary. Anything that can influence its context, including user messages, retrieved documents, tool results, and web pages, can influence its output. So whatever the model emits has the same trust level as the least trusted text it has read.

Once you accept that, the design follows. You wouldn’t let a REST API endpoint take a user_id from the request body and act on it without checking the session. Tool calls get the same treatment.

The validation layer for LLM tool call security

Our executor has one entry point, and every tool call goes through it. Simplified, it looks like this:

def execute(call, session):
    spec = TOOLS.get(call.name)
    if spec is None or call.name not in session.enabled_tools:
        return reject(call, "unknown_tool")

    args = spec.Args.model_validate_json(call.arguments)  # extra='forbid'
    spec.check_allowlist(args, session)
    spec.authorize(args, actor=session.user)

    return spec.handler(args, actor=session.user)

Each line covers a failure mode we’ve actually run into.

1. Tool registry scoped per session

The set of tools available is decided by the server for each session, not by whatever the model was told in its prompt. A support-chat session and an admin dashboard assistant get different registries. If the model emits a call to a tool that isn’t registered for this session, maybe because an injected document described a tool that exists elsewhere in the codebase, the call is rejected before any parsing happens.

2. Strict schemas, with extra fields rejected

Arguments are parsed into typed models with unknown keys forbidden (extra='forbid' in Pydantic, additionalProperties: false in JSON Schema, .strict() in Zod). This is the rule people push back on, because models are often helpful in ways nobody asked for. You define {order_id, amount, reason} and get back:

{"order_id": "A-1042", "amount": 40, "reason": "damaged",
 "notify_customer": true, "override_limit": true}

Dropping the unknown keys quietly seems harmless, until someone refactors a handler to **args into an ORM update and you’ve built mass assignment through a chatbot. Reject the call, return a structured error saying which field wasn’t allowed, and let the model retry. In practice the retry almost always comes back clean.

Provider-side “strict mode” for structured outputs helps, but only with shape. It guarantees the JSON parses against your schema. It says nothing about whether A-1042 belongs to the person in the chat.

3. Allowlisted argument values

Types aren’t enough. An amount is a number, but it has to fall between zero and the order’s refundable balance. A status is a string, but only certain transitions are legal from the current state. So we constrain values against server-side data. Enums for state changes, ranges computed from the database, and identifiers resolved through a query scoped to the session rather than a lookup by raw ID. If a scoped query can’t find the order, it doesn’t exist as far as this tool call is concerned.

4. Server-side authorization, re-checked

Identity never comes from arguments. The actor is always session.user, passed separately from anything the model produced. authorize() runs the same policy check the regular REST endpoint runs, ideally by calling the same function. If the tool path and the HTTP path have separate authz logic, they will drift apart, and the tool path is the one nobody pen-tests.

For destructive or irreversible actions, the handler doesn’t execute. It returns a pending action that a human confirms in the UI. The model can propose. The user signs off.

Why prompt-level defenses aren’t enough

Prompt-level mitigations include delimiting untrusted content, telling the model to ignore instructions inside documents, and running a classifier over inputs. They’re worth doing because they lower the rate of bad tool calls. But they’re probabilistic defenses against an adversary who gets unlimited retries and can see what works. A rule that holds 99% of the time is a good rule for tone. It isn’t a security control.

The deeper problem is where these defenses sit. They try to make the model trustworthy. The executor layer assumes it isn’t, and limits the damage to what the authenticated user was already allowed to do. With that layer in place, a fully compromised model can do at most what a malicious user with the same session could do through your normal UI. Without it, the model’s permissions are the union of every handler’s assumptions, and nobody has written that set down.

Two operational details make this layer hold up over time:

  • Log every rejection with the triggering context. Rejections are your prompt-injection telemetry. A spike in unknown_tool or forbidden-field errors usually points to a specific document or input pattern.
  • Return errors the model can act on, never stack traces. “Field override_limit is not permitted” gets a clean retry. A traceback leaks internals into a context window that someone else may be able to read.

None of this is new. It’s the same input validation, mass-assignment protection, and authz discipline backend developers have applied to web forms for twenty years. The only new part is remembering that the form is now being filled out by something that reads untrusted text for a living.

At orithLabs, this executor-first approach is how we build AI features into production backends, and it’s one of the first things we check when auditing an existing LLM integration. If you’re wiring tool calls into a system that matters and want a second pair of engineering eyes on the trust boundary, we’re happy to talk through it.