Three hikers asked Google Gemini to help plan meals for a Mount Shasta ascent. The plan it gave them skipped fat almost entirely, leaning hard on carbs and protein. On a multi-day climb above 10,000 feet, that’s the kind of gap that turns into hypoglycemia, poor decision-making, and slowed pace. According to the Siskiyou County Sheriff’s Office and reporting from Fox News, ABC News, and TechRadar, the group ended up stranded overnight and needed a rescue. Nobody died. But it’s exactly the kind of near-miss that should make anyone shipping an AI feature stop and ask a harder question than “does the model work.”
Because here’s the thing: Gemini isn’t a nutrition engine, and it was never claimed to be one. The failure isn’t that a general-purpose LLM gave incomplete advice when asked a domain-specific, safety-critical question — that’s a known, well-documented failure mode. The failure is that the product surface around it let that advice reach a hiker planning a real ascent with zero friction, zero caveat, zero routing to something more authoritative. That’s not a model problem. That’s a shipping problem.
The gap between “the model can answer” and “the model should answer”
Every LLM will answer almost any question you put in front of it, confidently, in a well-formatted list, with a tone of authority that has nothing to do with whether the answer is correct. That’s the core design tension of building on top of these models: capability and correctness are not the same axis, and general-purpose chat interfaces make no attempt to separate them for the user.
A nutrition question for a high-altitude, multi-day physical effort sits in a category we’d flag immediately in any AI feature review: high stakes, domain-specific, and dependent on individual variables (body weight, conditioning, medical history, expected exertion, weather) that a chat prompt typically doesn’t surface. That combination is exactly where you need one of a few things before the advice ships to an end user:
- A domain-scoped model or retrieval layer trained or grounded on actual mountaineering/nutrition guidance, not a general chat model improvising from broad training data.
- An explicit disclaimer and hard redirect to a human expert or established protocol (e.g., pointing to established altitude nutrition guidelines) when the query pattern matches a safety-critical category.
- A confidence or category flag that changes the UI treatment — the same way a search engine treats “how to treat a snakebite” differently from “what’s a good pasta recipe.”
None of that is exotic. It’s the unglamorous, unpaid-attention-to layer between “the model produced tokens” and “a user acted on those tokens in the real world.” It’s also exactly the layer that gets skipped when a feature ships fast and the assumption is that a general-purpose assistant is fine for everything, because it answers everything.
What a real pre-launch review actually checks
This is the review we run before any AI feature we build goes near a real user, and it’s not complicated — it’s just deliberate:
- Classify the query space. What categories of questions will this feature actually get asked? Which of those are consequential if wrong — financial, medical, physical safety, legal? Mount Shasta nutrition planning falls squarely in that bucket; most consumer AI chat products don’t classify at all.
- Decide what the model is allowed to answer directly. For high-stakes categories, the default should be narrower than “let the base model respond,” not wider. That might mean routing to a curated dataset, a rules-based fallback, or simply a “talk to a professional” deflection.
- Test with adversarial and naive prompts, not just happy-path ones. A user planning a first Shasta summit isn’t going to ask a hedge-covered, precisely-worded question. They’re going to ask something plain and underspecified — which is exactly the prompt shape that got these hikers a carb-only meal plan.
- Put a human in the loop somewhere before deployment, not just in the incident postmortem. If nobody on the team looked at output for the “plan my mountain climb” category before it went live, that’s the gap — not the transformer architecture underneath.
None of this is about mistrusting AI models. It’s about recognizing that a model is a component, not a product, and the product is where the responsibility for real-world consequences actually sits.
Security-mindedness applies to bad advice, not just bad actors
Most teams that think about AI security are thinking about prompt injection, data leakage, or jailbreaks — someone trying to make the model misbehave on purpose. That’s real and worth guarding against. But the Shasta case is a reminder that the more common failure won’t be adversarial at all. It’ll be an ordinary user asking an ordinary question in good faith, and the system handing back an answer with more confidence than the underlying model has any right to claim. Guarding against that isn’t a security feature you bolt on later — it’s a design decision you make before the feature ships, the same way you’d decide what happens when a payment fails or a form gets submitted twice.
We build AI features into production apps with that lens by default — treating “what happens when this is wrong” as part of the spec, not an afterthought discovered after a user gets stranded overnight. If you’re integrating an LLM into something people will actually rely on and want a second set of eyes on where the guardrails need to go, that’s a conversation we’re always up for.