Most founders we talk to can tell us what their AI feature will cost to build. Very few can tell us what it will cost to run in its third month, when a few hundred people are using it every day. The first number is a quote you pay once. The second is a bill that grows with your success, and that’s the number that decides whether the feature makes money. The good news is that LLM running costs are easy to estimate before anyone writes code. You need about six numbers and a spreadsheet.

Why the build quote is the wrong number to anchor on

A normal software feature costs almost nothing per use once it ships. A search bar or an admin dashboard runs on servers you’re already paying for. An LLM feature is different: every time a user presses “send”, you pay the model provider for the text going in and the text coming out, measured in tokens (roughly three-quarters of an English word each).

So an AI feature acts less like a one-off purchase and more like a new hire whose pay depends on how busy they are. That’s fine, as long as you know the rate before you commit. The worst version is the one where usage takes off, the invoice grows faster than revenue, and you’re rewriting the feature under pressure to make it cheaper.

A simple model for estimating LLM running costs

Here’s the per-user model we use at the planning stage. The numbers below are made up to show the method. Swap in your provider’s current price sheet and your own usage guesses.

  1. Input tokens per request. This means everything you send the model, not just what the user typed. Picture an in-app coaching assistant: a 1,500-token system prompt (instructions, tone, rules), about 2,000 tokens of recent conversation history, and a 200-token user message. That’s roughly 3,700 input tokens.
  2. Output tokens per request. Say 400 tokens for a useful reply.
  3. Price per token. For the example, assume $3 per million input tokens and $15 per million output tokens. Output almost always costs more than input, sometimes several times more.
  4. Cost per request. 3,700 × $3/1M ≈ $0.011, plus 400 × $15/1M = $0.006. Call it $0.017 per request.
  5. Requests per user per month. A typical active user might send 40 messages a month.
  6. Overhead multiplier. This covers everything that isn’t the main call. More on it below. Use 1.3 as a starting point.

That gives $0.017 × 40 × 1.3 ≈ $0.88 per typical user per month.

Don’t stop at the average user, though. Usage is almost never even. If 10% of your users are heavy users sending 300 messages a month, each one costs about $6.60. The blended cost is then (0.9 × $0.88) + (0.1 × $6.60) ≈ $1.45 per active user. That’s 65% higher than the “typical user” figure, and it’s the one you should compare against your pricing.

The hidden multipliers: retries, moderation, and growing context

The overhead multiplier is where estimates usually go wrong. It covers several kinds of cost that never show up in a demo.

  • Retries. Requests time out, return badly formatted responses, or hit rate limits. A production system retries some of them, and you pay for each attempt.
  • Moderation and safety checks. If users can type free text, you should be screening both what goes into the model and what comes back. We treat this as non-negotiable from an AI security standpoint. Those checks are usually cheaper model calls, but they happen on every request.
  • Multi-step calls. Features that classify the request first, fetch data, and then answer make two or three model calls per user action, not one.
  • Growing context. This is the one that surprises people. If you send the full chat history every time, message 20 costs far more than message 1, because the input keeps getting longer. Twenty turns in, that 2,000-token history can easily be 8,000 tokens.
  • Usage growth. People who find a feature useful use it more over time. Plan for per-user usage to rise over time, not stay flat.

A 1.3 multiplier is a reasonable first guess for a simple chat feature. An agent-style feature that calls tools and chains steps together can land well above 2.

Deciding whether the feature pays for itself

With a blended per-user cost in hand, the business question gets concrete. If a paid tier brings in a few dollars per user per month and the AI feature costs $1.45 of that, is the margin left after hosting, payment fees and support still acceptable? If the feature sits in a free tier, how many free users can one paying customer cover?

If the numbers look tight, that doesn’t mean you drop the feature. It means you design for cost from the start. These are the levers we use most:

  • Trim the context. Summarize older conversation instead of resending all of it, and keep system prompts short.
  • Use the right model for each step. Classification, routing and moderation rarely need the most capable (and most expensive) model.
  • Cache what repeats. Many providers discount repeated prompt prefixes, and common questions can be answered from your own cache.
  • Cap output length so the model can’t write an essay when a paragraph will do.
  • Set usage limits per plan. Heavy users can move to a higher tier instead of eating your margin.
  • Log tokens per request from day one, so your estimate becomes a measured number within weeks of launch.

None of these is exotic. They’re backend decisions that are cheap to make before launch and expensive to add later, which is why the cost model belongs in the planning conversation, not the post-launch cleanup.

When orithLabs scopes an AI integration, we go through this with you before quoting the build. We look at what each request actually sends, where the extra calls come from, and what the feature is likely to cost per user at the usage levels you expect. Sometimes the answer is to build it as planned, sometimes it’s a leaner version, and occasionally it’s “not yet”. If you’re weighing an AI feature and want a second pair of eyes on the running-cost side, we’re happy to go through your numbers with you.