Somewhere this week, a Salesforce admin is going to flip a toggle in Agentforce, watch Claude summarize a support case inside Slack, and assume the hard part is done. It isn’t. The hard part starts the moment that same admin asks the same agent to look up a customer’s billing history, and someone on the security team realizes an LLM now has a standing credential into a production data store.
Salesforce and Anthropic’s ‘Claudeforce’ push — Claude wired into Agentforce, Slack, and Salesforce’s developer tooling as a default reasoning layer — is a genuinely big deal, and not just as a press release. Slack alone sits on years of internal conversation, file shares, and DMs across thousands of companies. Making that corpus queryable by a model that can also take actions in Salesforce is a different risk category than a chatbot that answers questions about a product manual. We build production software for a living, not slideware, and the gap between “Claude is now embedded across the suite” and “this is safe and fast enough to trust with real customer data” is exactly the gap we get hired to close.
The keynote skips the auth boundary
Every LLM-in-a-SaaS-suite announcement glosses over the same question: what can the model actually reach, and on whose behalf? Agentforce agents acting inside Salesforce, pulling context from Slack, and touching custom objects means you’re composing three permission systems that were never designed together. Salesforce’s object-level security, Slack’s channel and workspace scoping, and whatever the model’s own tool-use layer thinks its access looks like — those three things drifting out of sync is not a hypothetical, it’s the default state until someone deliberately reconciles them.
On our own work — building AI features into Flutter apps and backend consoles, not enterprise CRMs — the pattern is identical at smaller scale. The moment you give a model tool access to a database or an internal API, you’ve created a new attack surface that doesn’t look like a traditional injection point, because the “attacker” input can just be a user typing a normal sentence into a support chat. Prompt injection via a Slack message, or via a case description a customer wrote, is a live path into whatever the agent is authorized to touch downstream. The fix isn’t “trust the model less” as a slogan — it’s concrete scoping work: least-privilege service accounts per agent action, explicit allowlists of which objects an agent can read versus write, and treating every piece of retrieved context as untrusted input even when it comes from inside your own Slack workspace.
Latency and fallback logic are the unglamorous 90%
The demo path for Claudeforce is a single clean request-response. Production usage is a support rep with forty tabs open, hitting an agent that has to call Salesforce’s API, maybe query Slack search, reason over both, and return an answer before the rep gives up and does it manually. Every one of those calls is a network hop with its own latency distribution, and Anthropic’s own current inference times mean a multi-tool-call agent response can easily run several seconds — fine for an async Slack thread, much less fine for a live console where someone’s waiting on a page to load.
What doesn’t make the announcement:
- What happens when the model call times out mid-workflow — does the agent retry, degrade to a cached answer, or fail the whole user action?
- How do you version-control prompts and tool schemas the same way you version-control code, so a prompt change doesn’t silently break a downstream integration?
- What’s the fallback when Claude is confident but wrong about which Salesforce record it just modified — is there a dry-run or confirmation step before a write action fires?
We hit smaller versions of all three building Crumb Count and Kompete. An AI feature that works in a demo and then falls over the first time the network hiccups isn’t a shipped feature, it’s a liability with good marketing. The unglamorous work — timeout handling, graceful degradation, structured logging so you can actually debug what the model decided and why — is what separates a Salesforce-blog screenshot from something a support team relies on at 2pm on a Tuesday during an outage.
What we’d actually check before rolling this out
If a company came to us asking to wire Claude (or any frontier model) into their existing product suite the way Salesforce is doing at massive scale, the audit questions are the same regardless of size:
- Map every data source the agent can touch and confirm scoping is enforced at the API layer, not just prompted for.
- Define what “the model is wrong” looks like for write actions specifically, and require confirmation or reversibility there.
- Load-test the full tool-calling chain, not just the model call, since the slowest hop sets your real latency.
- Decide up front what happens on model failure — silent fallback, visible error, or human handoff — instead of discovering the answer in production.
None of this is a reason to avoid embedding reasoning models into real products. It’s a reason to budget for the integration work as its own project, with its own security review, rather than treating it as a checkbox next to “enable Agentforce.”
This is the part of AI feature work we actually spend our time on — not the model choice, but the auth boundaries, the fallback paths, and the honest answer to “what happens when this breaks.” If you’re looking at wiring a reasoning model into an existing product and want a second opinion on where the real risk sits, that’s a conversation we’re glad to have.