Key takeaways
- Separate the model provider, the prompts, the retrieval, the business logic and the evaluation into modules that can change independently.
- Prompts are code. Store them as files, version them, and test them with a golden set.
- Validate every model output with a schema at the boundary; the rest of the code should never see raw model text.
- An eval command that runs a labelled set and prints a score is the single most valuable file in the repository.
- Typed settings, structured logs and a Makefile or task runner turn a notebook into something a second person can run.
How should I structure an LLM project in Python?
As a proper package under src/, with separate modules for the model provider, prompts, retrieval, domain logic, API or CLI, and evaluation; typed settings loaded from environment variables; prompts stored as versioned files rather than string literals; Pydantic schemas at every model boundary; a tests/ folder with a golden set and an eval command; and a task runner so make eval or uv run eval works on any machine. The structure exists so the model can change, the prompts can change and the retrieval can change without touching each other.
Also asked as: how to structure an llm project · llm project structure python · folder structure for llm application · genai project structure · rag project structure · python project structure for ai · llm app architecture python · how to organize an ai project · best project structure for langchain app · llm project template
Both production systems I maintain, the RAG platform and the document extraction service, share this shape, and the shape is why a provider swap in one took an afternoon rather than a week [11][12]. My shorter note on the layout is on the portfolio [13]. This page is the full version with the reasons.
The first time you change models, the structure pays for itself. The second time you change prompts without a test set, you learn why the eval folder exists. Pranjul Rathour
What does the folder layout look like?
A src/yourapp/ package with providers/, prompts/, retrieval/, core/, api/ and eval/ inside it; tests/ beside src/ with fixtures and a golden set; data/ for local indexes and samples, gitignored where large; scripts/ for one-off jobs; pyproject.toml, Makefile or task file, .env.example, README.md. The src layout prevents the common bug of importing the local folder instead of the installed package [1] and forces you to install the project properly.
Also asked as: llm project folder structure · src layout python · python ai project directory structure · rag app folder structure · where to put prompts in a python project · project layout for fastapi llm app · genai repo structure · langchain project folder structure
How should I handle configuration and secrets?
A single typed settings object loaded from environment variables, using Pydantic Settings or equivalent [2]: model names, endpoints, temperature, chunk sizes, feature flags, all with types and defaults. Secrets come from the environment and never from the repository; commit a .env.example with the keys and no values [10]. Every module receives the settings object or the specific values it needs; nothing reads os.environ directly. This is what lets the same code run on your laptop, in CI and in production.
Also asked as: how to manage config in python llm app · pydantic settings llm · environment variables for api keys python · where to store api keys in a python project · llm app configuration best practices · dotenv python llm · secrets management genai app · 12 factor config python
Where should prompts live?
In files, one per prompt, under src/app/prompts/, loaded by a small function that fills placeholders and records which version was used. A prompt is a piece of program logic that happens to be in English; treating it as a string literal buried in a function makes it impossible to diff, review or test. Name files by purpose and version, keep a changelog line at the top, and run the golden set whenever one changes. Templates with explicit placeholders beat f-strings, because the template can be reviewed without the surrounding code.
Also asked as: where to store prompts in code · prompt management python · prompt versioning · prompts as files · how to organize prompts in an llm project · prompt templates python · prompt engineering project structure · should prompts be in code or config
How do I handle model outputs safely?
Define a Pydantic model for every structured output and parse at the boundary [3]. Use the provider's structured output or tool-calling mode to constrain generation where available [7][8], then still validate, because constraints reduce but do not eliminate malformed output. On failure, retry once with the validation error in the prompt, then fall back to a safe default or a human queue. The rest of the codebase receives typed objects and never raw model text.
Also asked as: pydantic llm output · structured output python llm · validate llm output · llm json parsing python · how to handle llm response errors · instructor python · llm output schema · tool calling structured output python
DocuLens is built on this rule: a strict schema per document type, validation at the boundary, deterministic rules afterwards, and a review queue for anything that fails [12]. It is the difference between a demo and a system a finance team trusts.

How do I set up evaluation in the project?
A golden/ folder with labelled cases as JSON or JSONL: input, expected output or expected properties, tags. An eval module that runs every case through the real pipeline, computes metrics, retrieval recall, format validity, faithfulness judged by a model or rules, latency, cost, and writes a report. A command, make eval or uv run eval, that prints the score and exits non-zero if it drops below a threshold. Run it in CI on every prompt or retrieval change. Twenty cases is enough to start; grow it from production failures.
Also asked as: how to evaluate llm app · llm evaluation python project · golden dataset llm · eval folder structure · how to test llm outputs · llm evals in ci · rag evaluation script · regression testing prompts
How should I log and trace?
Structured logs, JSON in production, with a request ID that follows the call through retrieval, generation and validation [9]. Log the prompt version, model name, token counts, latency, retrieved document IDs and validation outcome for every call. Store prompts and outputs for a bounded time, with sensitive data redacted, so a failure report can be reproduced. When a user says the answer was wrong, you should be able to find exactly what the model saw.
Also asked as: logging for llm applications · llm observability python · how to debug llm app · trace llm calls · structlog llm · log prompts and responses · llm app monitoring · request id logging fastapi
Which tools should I use for the project itself?
uv for environments and dependencies, because it is fast and produces a lockfile [4]. Ruff for linting and formatting in one tool [5]. pytest for tests [6]. Pydantic for schemas and settings [2][3]. FastAPI if the app serves HTTP. A Makefile or a task runner with setup, test, eval, run and lint targets so nobody memorises commands. Pre-commit hooks that run Ruff. Nothing exotic; boring tools mean a new contributor is productive in an hour.
Also asked as: best python tools for ai projects · uv vs poetry · ruff python · python project setup 2026 · pyproject.toml llm project · python toolchain for genai · makefile python project · pre-commit python
How do I package and deploy it?
A Dockerfile that installs from the lockfile, copies src/, and runs the API with a production server. Settings from environment variables, so the same image runs everywhere. Health and readiness endpoints. Indexes and data mounted or fetched at startup, not baked into the image if they change. A CI pipeline that runs lint, tests and evals before building the image. Deployment specifics depend on the host; the structure makes them small.
Also asked as: deploy python llm app · dockerfile for llm application · fastapi llm deployment · how to deploy rag app · ci cd for llm project · containerize genai app · production python ai app
What mistakes make LLM projects unmaintainable?
Prompts as f-strings scattered through functions. Provider calls made directly from business logic, so a model swap touches everything. Raw model text parsed with regex in five places. No golden set, so every change is a guess. API keys in the repository. A notebook as the only entry point. No request IDs, so failures cannot be traced. Everything in one file called main.py at two thousand lines. Each is fixed by one folder on this page.
Also asked as: llm project mistakes · genai codebase anti patterns · why llm apps become unmaintainable · notebook to production python · refactor llm prototype · technical debt ai projects · common mistakes in rag codebases
Interview questions on LLM project structure
Explain how you would structure a RAG application so the model can be swapped. Where do prompts live and how are they versioned? How do you validate outputs? What is in your eval set and when does it run? How do you handle secrets? How would you debug a wrong answer reported by a user? A repository that answers these by its layout is the strongest evidence you can bring.
Also asked as: llm engineering interview questions · genai system design interview · how to design an llm application interview · rag system design interview · ai engineer code structure interview
Where should I start?
Take your current notebook and create the six folders. Move the model call into providers, the prompt into a file, the output into a schema, and five examples into a golden set with an eval command. Two hours, and the project can now change safely. For a workshop on shipping LLM applications from notebook to production, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.
Sources
- Python Packaging User Guide, src layout vs flat layoutpackaging.python.org
- Pydantic Settings documentationdocs.pydantic.dev
- Pydantic models documentationdocs.pydantic.dev
- uv, Python project managerdocs.astral.sh
- Ruff, Python linter and formatterdocs.astral.sh
- pytest documentationdocs.pytest.org
- OpenAI structured outputs guideplatform.openai.com
- Anthropic tool use documentationdocs.anthropic.com
- structlog documentationstructlog.org
- The Twelve-Factor App, config12factor.net
- RAG.NextUpgrad source code, Pranjul Rathourgithub.com
- DocuLens AI source code, Pranjul Rathourgithub.com
- Notebook to production checklist, Pranjul Rathourpranjulrathour.scult.in




