Key takeaways
- An AI app is an ordinary web app with a model call inside. The engineering that makes it survive is the ordinary part: API, database, auth, logs, tests, deploy.
- Python with FastAPI on the back, any simple front end, Postgres with pgvector, Docker, and one platform host is a stack that ships in a weekend and scales far.
- Free tiers exist for hosting, databases and model APIs, with limits. Design for those limits from day one and the app costs nothing until it has users.
- Every model call needs a timeout, a retry, a fallback, a token budget and a log line. Skip one and the first outage or bill will find it.
- Ship the smallest useful version to real people, log what they do, and iterate. A live app with ten users beats a perfect repository nobody runs.
How do I build an AI app?
Build a normal web application and put the model call behind one function. The front end collects input and shows output, the API validates the input, calls the model with a prompt and any retrieved context, validates the output, stores what happened, and returns it. Add authentication, a database, logging and a deploy pipeline, and you have an AI app. The model is one dependency, and everything around it is what makes the product work when a stranger uses it.
Also asked as: how to build an ai app · how to build ai application · how to create an ai app · ai app development · how to make an app with ai · how to build a chatgpt app · how to build llm application · build ai app from scratch · how to build an ai product
I shipped five production AI applications between February and July 2026, a RAG platform, a fine-tuning platform, a document extraction service, a browser face recognition system and an OCR and speech workspace, each with Docker, CI/CD, structured logging and metrics [12]. The stack below is the one they share, and the checklist at the end is what each of them taught me the hard way.
The model call is thirty lines. The other three thousand lines are why the app is still up next month. Pranjul Rathour
What is the architecture of an LLM app?
Five layers. A client: a web page, a mobile app, a WhatsApp bot. An API layer that authenticates, validates and rate-limits. A model layer: the prompt, the provider client, retries and fallbacks, and optionally retrieval, tools or a fine-tuned model. A data layer: a relational database for users and logs, a vector index for documents, object storage for files. And an operations layer: deploy pipeline, logs, metrics, alerts. Most AI apps are small enough that the API and model layers live in one service.
Also asked as: llm app architecture · ai app architecture · genai application architecture · rag application architecture · how to design an ai application · ai system design · llm application stack · ai backend architecture
The longer write-up of the FastAPI and Next.js version of this, with the folder layout, is on the portfolio [13].
Which tech stack should I use for an AI app?
Python for the back end, because every model library, SDK and example is in Python, with FastAPI for the API [1]. Postgres for data, with pgvector when you need embeddings [3]; Supabase gives you Postgres, auth and storage on a free tier [4]. Any front end you already know, plain HTML, React or Next.js. Docker to package it [2]. GitHub Actions to test and deploy [8]. One platform host. Add nothing else until a real need appears.
Also asked as: best tech stack for ai app · tech stack for llm app · python or javascript for ai app · fastapi for ai · next js ai app · best framework for ai app · ai app tech stack 2026 · langchain vs plain python · full stack ai app
On frameworks such as LangChain: excellent for a first prototype and for loaders; for production, I write the model wrapper myself, because a bug under three abstraction layers costs a day and a bug in my own fifty lines costs ten minutes.
How do I deploy an AI app for free?
Package the app in Docker and deploy to a platform with a free or trial tier: Railway, Render or Fly.io for a container, Hugging Face Spaces for a demo with a GPU option, Supabase for the database, and a static host such as GitHub Pages, Netlify or Cloudflare Pages for the front end. Free tiers sleep or cap hours, so design the app to start fast and hold no state in memory. The model API itself is not free beyond trial credits, so budget tokens from day one.
Also asked as: how to deploy ai app for free · free hosting for python ai app · deploy fastapi for free · free hosting for llm app · host ai app free · railway vs render free tier · deploy python app free · where to host ai app · free backend hosting 2026 · hugging face spaces deploy
Free tiers change; the platform documentation is the source of truth and this table is a map, not a guarantee [4][5][6][7]. RAG.NextUpgrad, with hybrid retrieval, reranking and streaming, runs in about 220 MB of memory on a free tier, which is the proof that a real AI app does not need a large server [12]. My deployment write-up walks through it step by step [14].

How much does it cost to run an AI app?
Hosting can be zero at first. The model API is the real bill: input tokens plus output tokens, times the price per million, times requests. Estimate it before launch: measure the tokens of a typical request, multiply by expected daily requests, multiply by the provider's rate. Then set a hard monthly cap at the provider, a per-user daily limit in your code, and an alert at half the cap. An app with no cap is one viral post away from a bill you did not plan.
Also asked as: how much does it cost to run an ai app · ai app cost · llm api cost for app · cost of building ai app · how to reduce ai api costs · token budget · ai app pricing model · openai api cost per user
How do I make an AI app reliable?
Wrap every model call with a timeout, a bounded retry with backoff, and a fallback to a second provider or a smaller model with the same prompt. Validate outputs against a schema and fail closed when they do not parse. Stream responses so users see progress. Make the app stateless so it can restart. Test the prompt against a fixed set of inputs in CI. Log every request with its tokens, latency and outcome. My apps fall back across providers automatically, and that single feature has kept them up through several API outages.
Also asked as: ai app reliability · llm api fallback · handling llm api errors · llm timeout retry · multi provider llm · ai app error handling · how to handle openai downtime · graceful degradation ai app · llm rate limits handling
async def generate(prompt: str, schema: type[BaseModel]) -> BaseModel:
for provider in PROVIDERS: # identical prompt, ordered by preference
for attempt in range(2):
try:
raw = await asyncio.wait_for(provider.complete(prompt, max_tokens=800), timeout=30)
return schema.model_validate_json(raw) # fail closed if it does not parse
except (TimeoutError, ProviderError, ValidationError) as e:
log.warning("provider=%s attempt=%d err=%s", provider.name, attempt, e)
await asyncio.sleep(0.5 * (attempt + 1))
raise ServiceUnavailable("all providers failed")
How do I secure an AI app?
Treat every input, and every document or tool result the model reads, as untrusted. Keep instructions in the system prompt and data in delimited user content. Give tools the least permission they need and require a human confirmation for irreversible actions. Never put secrets in prompts. Rate-limit per user. Validate outputs before acting on them. Log prompts and outputs with user consent and a retention policy. OWASP's list of LLM application risks is the checklist to review before launch [9].
Also asked as: ai app security · llm security best practices · prompt injection prevention app · securing openai api key · api key in frontend · ai app privacy · owasp llm top 10 · how to protect ai app from abuse
The most common student mistake is the API key in the front-end code, where anyone can read it and run up your bill. The key lives on the server, always, and the server enforces the limits.
Should I build an AI agent or a simple app?
A simple app, until a decision appears that you cannot write as code. Most "AI agents" that succeed in production are workflows: code decides the steps, a model does the language work at each step. Reach for a real agent loop only when the task is open-ended and the model must choose actions at run time, and then add limits, permissions and confirmations before anything else [11].
Also asked as: ai agent vs ai app · should i build an ai agent · agentic app architecture · workflow vs agent · when to use ai agents in app · ai automation vs ai agent
How do I get users for an AI app?
Build it for a group you belong to, your class, your club, a family business, and put it in front of them the day it works. Ask what they tried to do that failed, and fix that. Post a demo video where those people are. Ten real users teach more than a thousand sign-ups from a launch post, and they are how a student project becomes a line on a resume that a recruiter believes.
Also asked as: how to get users for ai app · launch ai app · how to market an ai app · ai side project users · validate ai app idea · build in public ai app · ai app for students
What mistakes kill AI apps after demo day?
No cap on API spend. The key in the front end. No logs, so the first complaint is undebuggable. State held in memory on a host that sleeps. Prompts edited in production with no test set. One provider and no fallback. No per-user limits, so one script drains the budget. A model output acted on without validation. Each is on the checklist above; each I have either done or watched a mentee do.
Also asked as: ai app mistakes · why ai side projects fail · common llm app bugs · ai app production issues · ai project after hackathon · student ai project mistakes
What are common AI app interview questions?
Design an LLM feature end to end and name where it can fail. Estimate its monthly cost. Explain how you would handle a provider outage. Explain where prompt injection can enter and what limits its damage. Describe how you would test a prompt change. Say what you would log. Each answer is a section above, and the best evidence is a live app you can open during the interview.
Also asked as: ai app interview questions · llm system design interview · genai system design · ai engineer system design questions · how to explain ai project in interview
Where should I start?
Take a problem someone near you has, build the smallest version with FastAPI, one model call with a timeout and a fallback, Postgres on Supabase, Docker, and deploy it on a free tier this weekend. Log everything. Show it to three people. Then read the checklist again. All five of my production apps are open source and follow this shape [12]. If your college wants a workshop that ends with every student having a deployed AI app, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.
Sources
- FastAPI documentationfastapi.tiangolo.com
- Docker documentationdocs.docker.com
- pgvector: vector similarity search for Postgresgithub.com
- Supabase documentation (Postgres, auth, storage)supabase.com
- Railway documentationdocs.railway.com
- Render documentationrender.com
- Hugging Face Spaces documentationhuggingface.co
- GitHub Actions documentationdocs.github.com
- OWASP Top 10 for Large Language Model Applicationsowasp.org
- Twelve-Factor App12factor.net
- Anthropic, Building effective agents (2024)anthropic.com
- RAG.NextUpgrad source code, runs in about 220 MB on a free tier, Pranjul Rathourgithub.com
- FastAPI and Next.js AI app architecture, Pranjul Rathourpranjulrathour.scult.in
- Deploying a Python AI app for free, Pranjul Rathourpranjulrathour.scult.in




