How much does an LLM API cost? Token pricing explained, a cost model you can run, and 12 ways to cut the bill without cutting quality

LLM APIs charge per million tokens, input and output priced separately. How to estimate the monthly cost of a feature before you build it, why RAG prompts are expensive, what prompt caching, batching, routing and smaller models save, when self-hosting an open model is cheaper, and the caps and alerts that stop a viral day from becoming a bill. From running five production AI apps on tight budgets.

Walking a room through the products he has shipped
Walking a room through the products he has shipped

Key takeaways

  • Cost per request is input tokens times the input rate plus output tokens times the output rate. Output usually costs several times more per token.
  • Measure a real request's tokens, multiply by expected daily volume, and you have a monthly estimate before writing production code.
  • The biggest savings are structural: send fewer passages, cache repeated prefixes, route easy tasks to small models, cap output length.
  • Self-hosting an open model wins at high steady volume or strict privacy, and loses at low or spiky volume once you count engineering time.
  • Hard spending caps, per-user limits and alerts are not optional. Set them before the first external user.

How much does an LLM API cost?

LLM APIs charge per token, priced per million tokens, with input, what you send, and output, what the model writes, priced separately, and output typically costing three to five times more per token. Your cost for one request is input tokens times the input rate plus output tokens times the output rate. Your monthly cost is that times the number of requests. Rates differ by model tier by more than a hundredfold, and they change often, so the providers' pricing pages are the only source of current numbers [1][2][3].

Also asked as: how much does an llm api cost · llm api pricing · llm cost · cost of using chatgpt api · openai api cost · claude api cost · gemini api cost · llm pricing per token · how much do llms cost to run · ai api cost

I will not print price figures on this page, because they would be wrong within months and this site only states what can be checked. What does not change is the arithmetic, and the arithmetic is what lets you estimate a feature's cost before you build it. My five production AI apps in 2026 all ran on small budgets; one, a full RAG platform, fits in about 220 MB on a free tier [13]. The habits that made that possible are below, and the longer notes are on the portfolio [10][11].

The API bill is a design decision you made in the prompt template. Every passage you send, every token you let the model write, is a line item. Pranjul Rathour

What is a token, and how many are in my request?

A token is the unit the model reads and writes, roughly three-quarters of an English word, more for code and for Indic scripts. A request's input tokens include the system prompt, the conversation history, any retrieved documents and the user's message; output tokens are the answer. Count them with the provider's tokeniser before you estimate anything [9]. A "short question" over a RAG system can be three thousand input tokens once five passages are attached.

Also asked as: what is a token in llm pricing · how many tokens in a word · token counter · how to count tokens · input tokens vs output tokens · tokens per request · token cost calculator · llm token limit

How do I estimate the monthly cost of an AI feature?

Build one real request, count its input and output tokens, and multiply. Then multiply by requests per day and by thirty. Then multiply by two, because production prompts grow and users retry. Do this before building, and again after the first week of real traffic, because estimates are usually wrong in the same direction: more tokens per request than planned, and more requests from a few heavy users than from the average.

Also asked as: how to estimate llm cost · llm cost estimation · ai feature cost calculator · cost per query llm · monthly cost of chatbot · openai cost per user · how to calculate api cost · llm budget planning

def monthly_cost(in_tok: int, out_tok: int, in_rate: float, out_rate: float, per_day: int) -> float:
    """in_rate/out_rate are the provider's prices per million tokens for the chosen model."""
    per_request = (in_tok * in_rate + out_tok * out_rate) / 1_000_000
    return per_request * per_day * 30 * 2     # x2: growth, retries, prompt creep

Why are RAG systems expensive, and how do I make them cheaper?

Because every question resends the retrieved passages, and passages are the largest part of the prompt. A system that stuffs ten chunks pays for ten chunks on every turn. Fixes, in order of impact: rerank to five or fewer passages; keep chunks tight and structure-aware so each passage carries only what is needed; cache the system prompt and any fixed prefixes where the provider supports it [4]; trim conversation history to the last few turns plus a summary; and refuse early, with a confidence gate, so unanswerable questions never reach the model.

Also asked as: rag cost optimization · why is rag expensive · reduce rag token usage · rag prompt cost · how many chunks to send to llm · context window cost · rag latency and cost · cheap rag

Sending fewer, better passages also improves answers, not just cost: models attend poorly to the middle of long contexts [8]. In RAG.NextUpgrad, reranking to five passages and gating weak queries cut prompt size and stopped the most expensive category of request, the confident wrong answer that a user then retries three times [13].

What is prompt caching, and how much does it save?

Prompt caching lets a provider store a prefix of your prompt, typically the system prompt, tool definitions and large fixed documents, and charge a much lower rate when the same prefix is reused within a time window [4]. For applications with long system prompts or shared documents across many requests, the saving on input tokens is large. It requires structuring prompts so the fixed part comes first and the variable part last, which is good prompt hygiene anyway.

Also asked as: prompt caching · what is prompt caching · prompt caching cost savings · how to use prompt caching · cached tokens pricing · context caching gemini · reduce input token cost

How do I cut LLM costs without hurting quality?

Twelve levers, roughly in order of payoff. Most teams pull three and stop; the first six are structural and the rest are operational.

Also asked as: how to reduce llm cost · llm cost optimization · reduce openai api cost · cut ai api bill · llm cost saving techniques · optimize token usage · cheaper llm alternatives · llm cost best practices

Is self-hosting an open model cheaper than an API?

At high, steady volume or under strict privacy requirements, yes: a GPU serving an open model through vLLM has a fixed monthly cost regardless of tokens [7]. At low or spiky volume, no: the GPU sits idle, and you pay for the engineer who keeps it running. The break-even depends on your token volume, the model size you need, and whether you already have operations capacity. Many teams do both: a small self-hosted model for cheap high-volume tasks and an API for the hard ones.

Also asked as: self hosting llm cost · api vs self hosted llm · is self hosting llm cheaper · cost of running llm on gpu · open source llm cost · vllm cost · llm hosting cost comparison · gpu cost for llm inference

Walking a room through the products he has shipped
Walking a room through the products he has shipped

Which provider is cheapest?

It changes monthly, and cheapest per token is the wrong question. Compare cost per solved task: a cheaper model that needs two retries or a longer prompt costs more. Use a routing layer such as OpenRouter to compare prices and models across providers without many keys [6], test the top two or three on your evaluation set, and pick by quality per rupee. My decision framework for choosing a provider, with the questions beyond price, is on the portfolio [12].

Also asked as: cheapest llm api · which llm api is cheapest · llm price comparison · openai vs anthropic vs gemini pricing · cheapest ai model api · llm pricing comparison 2026 · best value llm api

How do I stop a surprise bill?

Set a hard monthly spending cap at every provider. Add per-user daily request and token limits in your code. Set max output tokens on every call. Alert at fifty percent of budget and again at eighty. Log tokens and cost per request with the user id. Fall back to a smaller model or a polite refusal when the daily budget is exhausted. Never ship an API key in client code. One viral post or one abusive script is enough to find the control you skipped.

Also asked as: openai spending limit · how to set api budget · llm cost alerts · prevent api bill shock · api key leaked cost · rate limiting llm app · usage caps ai app · monitor llm spend

What does an LLM feature cost at student or startup scale?

Small, if designed with the levers above. A RAG assistant for a college club with a few hundred questions a day, five reranked passages, a capped answer and a mid-tier model costs less per month than a couple of coffees; the same assistant with ten unranked passages, unlimited output and a top-tier model for every request costs many times more for slightly worse answers. Free tiers from providers cover prototypes entirely. The difference is design, not budget.

Also asked as: cost of ai chatbot for small business · llm cost for startup · ai app cost for students · how much does it cost to run a chatbot · ai project budget · cheap ai for small business · free tier llm api

LLM cost interview questions

Estimate the monthly cost of a described feature out loud. Explain why output tokens matter more. Name three structural ways to cut a RAG system's cost. Explain prompt caching. Say when you would self-host and how you would decide. Describe the controls that prevent a runaway bill. Interviewers use these to find out whether you have run something real; a number from a system you operated is the best answer.

Also asked as: llm cost interview questions · ai system design cost estimation · genai interview cost optimization · token budgeting interview

Where should I start?

Take one request from your current or planned app, count its tokens with the provider's tokeniser, run the cost model above with today's prices, and set a spending cap before you do anything else. For a session on shipping AI features that survive real users and real budgets, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. OpenAI API pricingopenai.com
  2. Anthropic API pricinganthropic.com
  3. Google Gemini API pricingai.google.dev
  4. Anthropic prompt caching documentationdocs.anthropic.com
  5. OpenAI Batch API documentationplatform.openai.com
  6. OpenRouter: model pricing across providersopenrouter.ai
  7. vLLM documentation, high-throughput serving of open modelsdocs.vllm.ai
  8. Liu et al., Lost in the Middle: How Language Models Use Long Contexts (2023)arxiv.org
  9. tiktoken: OpenAI tokeniser for counting tokensgithub.com
  10. LLM cost control with token budgets, Pranjul Rathourpranjulrathour.scult.in
  11. Cost alerts for AI APIs before the bill surprises you, Pranjul Rathourpranjulrathour.scult.in
  12. Choosing an LLM provider: a decision framework, Pranjul Rathourpranjulrathour.scult.in
  13. RAG.NextUpgrad, runs in about 220 MB on a free tier, Pranjul Rathourgithub.com
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on Shipping AI apps

All Shipping AI apps guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur