Building a Chatbot API with FastAPI + OpenAI/Claude — Step by Step
Posted on Wed 30 September 2026 in Tutorial
Why This Matters
Every GenAI journey eventually hits the same milestone: building your first real chatbot API, not just calling a model from a script. This is the project that ties together everything else — API keys, request/response models, calling an LLM provider, and returning a clean response your frontend can actually use.
This walkthrough builds a simple but properly structured chatbot API using FastAPI, that you can extend later with streaming, memory, or RAG.
What We're Building
A minimal API with one core endpoint:
POST /chat → { "message": "your question" } → { "response": "the AI's answer" }
Simple on purpose. Once this works, adding conversation history, streaming, or tool use is much easier on top of a solid base.
Step 1: Project Setup
mkdir chatbot-api && cd chatbot-api
python -m venv venv
source venv/bin/activate # on Windows: venv\Scripts\activate
pip install fastapi uvicorn python-dotenv pydantic-settings
pip install openai # if using OpenAI
pip install anthropic # if using Claude
Project structure:
chatbot-api/
├── main.py
├── config.py
├── .env
├── .gitignore
Step 2: Managing the API Key Properly
As covered in an earlier post on environment variables, never hardcode API keys. Set up .env and .gitignore:
# .env
OPENAI_API_KEY=sk-your-real-key-here
# .gitignore
.env
venv/
# config.py
from pydantic_settings import BaseSettings, SettingsConfigDict
class Settings(BaseSettings):
openai_api_key: str
model_config = SettingsConfigDict(env_file=".env")
settings = Settings()
Step 3: Defining the Request and Response Models
FastAPI leans on Pydantic to define exactly what a valid request looks like:
# main.py
from pydantic import BaseModel
class ChatRequest(BaseModel):
message: str
class ChatResponse(BaseModel):
response: str
This gives you free input validation — if message is missing or the wrong type, FastAPI rejects the request with a clear error before your code even runs.
Step 4: Calling the LLM (OpenAI Version)
# main.py
from fastapi import FastAPI
from openai import OpenAI
from config import settings
app = FastAPI()
client = OpenAI(api_key=settings.openai_api_key)
@app.post("/chat", response_model=ChatResponse)
def chat(request: ChatRequest):
completion = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": request.message}
]
)
reply = completion.choices[0].message.content
return ChatResponse(response=reply)
Step 4 (Alternative): Calling the LLM (Claude Version)
If you're using Anthropic's API instead:
# main.py
from fastapi import FastAPI
from anthropic import Anthropic
from config import settings
app = FastAPI()
client = Anthropic(api_key=settings.anthropic_api_key)
@app.post("/chat", response_model=ChatResponse)
def chat(request: ChatRequest):
message = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1000,
messages=[
{"role": "user", "content": request.message}
]
)
reply = message.content[0].text
return ChatResponse(response=reply)
The overall shape is nearly identical between providers — build a request, send it, extract the text from the response. This is why swapping providers later isn't usually a huge rewrite.
Step 5: Running the API
uvicorn main:app --reload --port 8000
Test it directly from the auto-generated docs at http://localhost:8000/docs — no frontend needed yet. Send a request through the interactive UI and confirm you get a real response back from the model.
You can also test with curl:
curl -X POST http://localhost:8000/chat \
-H "Content-Type: application/json" \
-d '{"message": "What is FastAPI?"}'
Step 6: Handling Errors Gracefully
Right now, if the LLM API call fails (rate limit, network issue, invalid key), the error bubbles up as an ugly 500 response. Wrap it properly:
from fastapi import HTTPException
@app.post("/chat", response_model=ChatResponse)
def chat(request: ChatRequest):
try:
completion = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": request.message}
]
)
reply = completion.choices[0].message.content
return ChatResponse(response=reply)
except Exception as e:
raise HTTPException(status_code=502, detail=f"AI provider error: {str(e)}")
A 502 here signals to the client that the failure came from the upstream AI provider, not from your own API — a small detail that makes debugging much easier later.
Step 7: Adding Basic Conversation Memory (Optional Next Step)
The version above has no memory — every request is treated as a brand-new conversation. A simple way to add short-term memory is accepting the conversation history from the client and passing it straight through:
from typing import List
from pydantic import BaseModel
class Message(BaseModel):
role: str # "user" or "assistant"
content: str
class ChatRequest(BaseModel):
messages: List[Message]
@app.post("/chat", response_model=ChatResponse)
def chat(request: ChatRequest):
completion = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": m.role, "content": m.content} for m in request.messages]
)
reply = completion.choices[0].message.content
return ChatResponse(response=reply)
This pushes memory management to the client for now (it resends the full history each time) — a reasonable starting point before introducing server-side session storage or a database.
What to Add Next
Once this base is working, natural next steps (each a solid follow-up project) include:
- Streaming responses using Server-Sent Events, so replies appear word by word instead of all at once
- Rate limiting so one user can't drain your API budget
- Server-side conversation storage, instead of relying on the client to resend history every time
- RAG integration, grounding responses in your own documents instead of just the model's training data
Closing Thought
A chatbot API looks simple from the outside — send a message, get a response — but building it properly, with clean request validation, safe key management, and real error handling, sets up a foundation that's actually easy to extend later. Get this base right first, and streaming, memory, and RAG all become incremental additions instead of a rewrite.