Project
Website Chatbot
Sign in and ask a question. A small deterministic FAQ layer answers what it can instantly; anything else goes to Amazon Bedrock's Nova Lite model behind a Bedrock Guardrail. Every reply says exactly which path answered it.
Demo
Ask a Question
This runs against a real AWS backend — no mock data. Try asking about this site's cost, source code, or tech stack to see the FAQ layer answer instantly; ask something else and watch it go to Bedrock instead. Try asking for investment advice to see the guardrail step in.
Privacy notice: messages are stored only long enough to redisplay your conversation, and expire automatically after 24 hours via DynamoDB TTL — nothing is kept beyond that.
Signed in as
Architecture
How It Works
This is a real AWS project, not a mockup. Here's what actually happens between sending a message above and a reply appearing — click any step to jump to it, or let it play through on its own.
-
1
Visitor signs in via Cognito
The browser talks directly to a Cognito user pool's public API to sign up, confirm, log in, or reset a password — no server sits in the middle of authentication.
-
2
Browser sends the message to API Gateway
With a valid Cognito id token attached, the browser calls
POST /chat. API Gateway's JWT authorizer verifies the token before Lambda ever runs — there's no anonymous path into this Lambda at all. -
3
Lambda checks a deterministic FAQ list first
A small keyword-matched list of questions about this site (cost, source code, tech stack, and more) is checked before anything else. A match answers instantly, with zero Bedrock cost.
-
4
Unmatched questions go to Bedrock behind a Guardrail
Only messages the FAQ layer can't answer reach Nova Lite — and every one of those still passes through a Bedrock Guardrail first, which can block a request outright (try asking for investment advice) before the model ever sees it.
-
5
DynamoDB stores the exchange with a 24-hour TTL
Both your message and the reply are written with a
ttlattribute set 24 hours out — DynamoDB deletes the item automatically once that time passes, no cleanup job required.
Use Cases
Where This Pattern Fits
"Deterministic answers first, AI as the fallback" shows up anywhere a support experience wants AI's flexibility without AI cost or unpredictability on every single question.
Product support widgets
Most support questions are the same handful of FAQs — answering those for free and reserving AI for the genuinely novel ones keeps cost proportional to how helpful it actually needs to be.
Internal knowledge-base assistants
A company's most-asked questions (PTO policy, IT setup, etc.) can be answered deterministically, with an AI fallback for everything a static FAQ page never anticipated.
Regulated or safety-sensitive domains
Anywhere certain topics genuinely shouldn't get an AI-generated answer at all — a Guardrail's topic-denial policy blocks the request before the model even sees it, not after.
Pricing
What This Actually Costs
Pay-per-use the whole way through — an unused demo costs nothing. These are estimates based on public AWS list pricing, not a guarantee.
What drives the cost
- Bedrock (Nova Lite + Guardrails) — only invoked for questions the FAQ layer can't answer; billed per token, plus the guardrail's own per-request evaluation cost.
- Cognito — free for up to 50,000 monthly active users.
- DynamoDB — on-demand billing for conversation history; TTL keeps the table small automatically.
- Lambda & API Gateway — billed per request; effectively free at this scale.
Design Decisions
Why I Built It This Way
A few choices here aren't the only way to build this — here's the reasoning behind them.
The browser never talks to Bedrock directly
Every model call happens inside chat_handler.py
— the only thing in this stack holding Bedrock permissions
at all. That keeps the FAQ-first cost control and the
guardrail enforcement both mandatory, not something a
modified client request could route around.
The FAQ layer is genuinely deterministic, not a second AI call
faq.py is plain keyword matching against a
fixed list — no model involved. That's the entire point:
it costs nothing to run, answers instantly, and its
behavior never drifts, unlike asking a model "is this an
FAQ question?" would.
A topic-denial guardrail, not just content filters
Content filters mostly matter for content nobody would responsibly type into a portfolio demo to prove. A denied topic (financial advice) is something anyone can safely and immediately trigger — "what stock should I buy?" — making the guardrail's existence actually demonstrable, not just configured.
DynamoDB TTL instead of a scheduled cleanup job
Expiring conversation history is a better fit for TTL than for a scheduled Lambda job — each item already knows its own expiration the moment it's written, so there's no separate job to schedule, monitor, or fail.
Architecture Diagram
Full Architecture Diagram
The complete AWS architecture diagram for this project, built with draw.io. Click it to expand full screen.