Customer Support Agent
A tool-calling agent that answers order questions from the retailer's real customers about their real invoices. It looks things up, applies the refund policy in code, cites the policy, and stops for a person before any refund over £50. Every step is shown.
1Problem
Most customer messages to an online shop are about one order: something arrived broken, something is missing, a parcel is late. Answering them means opening the invoice, checking tracking, reading the policy and working out what is owed. It is repetitive, and mistakes cost money in both directions.
Chatbots that answer from a prompt alone make it worse: they promise refunds the policy does not allow, invent order details, and can be talked into things.
2Why it matters
An agent is only useful in operations if it is safe to let it act. That means its facts come from systems, its amounts from rules, its actions are limited and reviewable, and a person decides anything significant. Those are the properties this project demonstrates, with a trace that lets anyone check what happened.
3Solution
The agent works through tools. It reads the invoice and the tracking, searches the policy, and asks the refund rules what is due. It can record a refund only for the exact amount the rules returned: up to £50 it does so itself; above that, it asks for approval and the conversation waits for the supervisor. It opens tickets for work a person must do and drafts a reply that cites the policy. Nothing is sent and no money moves.
A free-tier model drives the loop when one is configured. Without one, a scripted agent takes the same steps with the same tools, so the demo always works and the difference between the two can be measured.
4Architecture
Customer message
About one of their real invoices
Agent loop
Up to 8 steps and 25 seconds a turn. Each step: the model sees the conversation and the tools, proposes calls; arguments are validated, tools run, results go back.
Provider chain
Groq, then Cerebras (free tiers, when keys are set), then the scripted agent. Token budgets per day and per visitor.
get_invoice
Real invoice lines and totals
list_customer_invoices
The customer's recent orders
get_shipment_status
Simulated tracking for the invoice
search_policy
Full-text search, numbered passages
calculate_refund
Policy rules in code: eligibility and amount
request_approval
Records a refund; over £50 waits for a supervisor
create_ticket
Hands work to a person
draft_reply
Saves the reply; nothing is sent
Real. Customers (anonymous IDs), invoices, products, quantities and prices from UCI Online Retail II.
Simulated. Shipments, the service policy, tickets, refunds and the messages themselves. Nothing leaves the system.
Human in the loop. A refund above £50 pauses the conversation until the supervisor (you) approves or rejects it.
5Technical implementation
- Agent loop
- Written by hand against the OpenAI-compatible chat API; no agent framework
- Models
- Groq and Cerebras free tiers when configured, a scripted agent as the last step
- Tools
- Eight Python functions with Pydantic argument models, scoped to one customer
- Data
- Real invoices in a read-only DuckDB snapshot; policy search in PostgreSQL
- API
- FastAPI with per-visitor quotas; conversations kept for 24 hours
- Interface
- Next.js: chat, supervisor approval and a step-by-step trace
The loop. Each step sends the conversation and the tool definitions to the provider chain. All tool calls in a response run as a batch (looking up the invoice and its tracking happens in one step), each result is returned to the model, and the loop ends when the agent drafts a reply or a refund needs approval. A supervisor's decision is added to the conversation and the loop resumes.
Refund rules. Damaged or missing items within 14 days of delivery, late deliveries (more than 3 days after the estimate: the shipping charge, or 5% of the goods up to £10), lost parcels (10 days past the estimate) and returns within 30 days (refunded on receipt). Credit notes are never refunded again.
Real and simulated data. Invoices come from 797,885 real invoice lines of customers with an ID, after removing the source export's duplicate lines. Shipments are simulated from each invoice's number and date, so the same invoice always has the same tracking; carrier services have generic names. The desk works as of 10 December 2011, the day after the data ends.
Policy search. The policy is split into passages by section and searched with PostgreSQL full-text search through the shared retrieval package; passages are numbered so replies can cite them, and a reply cannot cite a passage the agent did not retrieve.
6Live demonstration
Pick a scenario, or write your own message as one of the customers. The left side shows the conversation; the right side shows every step the agent took. When a refund needs approval, you decide.
Loading scenarios…
Portfolio demonstration. Customers and invoices are real (anonymised by the source); shipments, the policy, tickets and messages are simulated. Each visitor can start 20 conversations an hour.
7Results
Results are read from the running system and are not available right now. They appear after the first scheduled run of the scenario suite.
8Failure handling
- The model proposes invalid arguments or an unknown tool
- Arguments are validated before anything runs. The error goes back to the model as the tool result, so it can correct itself on the next step.
- The model, or the customer, tries to raise a refund
- Refund amounts come only from calculate_refund. request_approval refuses any amount that differs from the rules' result by more than a penny, and refuses to run before a calculation.
- A customer asks about someone else's order
- Every tool is scoped to the conversation's customer. Another customer's invoice gets the same answer as an invoice that does not exist, so nothing leaks.
- The agent loops or stalls
- A turn stops after 8 steps or 25 seconds. The conversation is then handed to a person with a ticket, and the customer is told so.
- No model is available, or the free tier is used up
- The provider chain falls through to the scripted agent, which uses the same tools and rules. The trace says which provider answered each step and why earlier ones were skipped.
- A large refund needs a decision
- Above £50 the agent can only request approval. The conversation pauses, the visitor decides as supervisor, and the agent drafts the follow-up from the decision.
- Anything would reach a real person or account
- Nothing can: refunds, tickets and replies are records in this demo's database, deleted after a day.
9Deployment
The agent runs in the site's API container. The order snapshot is built into the image from the verified dataset download; the policy is indexed into PostgreSQL the first time it is searched. Conversations are stored for 24 hours and removed by a daily job, which also re-runs the scripted baseline.
nginx gives the agent's routes a longer read timeout (a turn with a model can take several seconds) and a stricter request rate than the rest of the API.
10Cost
$0 in API spend. Only free tiers are ever used, through a token ledger with daily limits per provider and per visitor. When the limits are reached the chain falls through to the scripted agent, which costs nothing.
Every model step resends the instructions and the eight tool definitions, about 5,600 characters (roughly 1,400 tokens by the ledger's estimate of four characters per token), plus the conversation and tool results so far. The trace shows the measured tokens of each step when a model answers.
11Limitations
- The scripted agent reads keywords and product names. It handles the scenarios it was written for and simple variations; anything else gets a clarifying question or a hand-over.
- Models have not been scored on this server yet, because none is configured by default. The suite is ready to run when a key is.
- Shipments and the policy are simulated, and credit notes in the source data do not reference the invoice they credit, so the agent cannot check for an earlier refund on the same items.
- Each conversation is one customer; there is no login. The customer is chosen in the demo, which a real system would take from an authenticated session.
12Source code
The full source is on GitHub under the MIT licence, with the tests and scenarios.
- agent.pyThe agent loop, budgets, hand-over and approval resume
- tools.pyThe eight tools and their argument models
- refunds.pyRefund rules
- scripted.pyThe scripted agent used without a model
- shipments.pySimulated shipments
- policy.pyThe written policy and its indexing
- scenarios.pyThe customer scenarios
- evaluation.pyScenario checks
Data: Chen, D. (2012). Online Retail II [Dataset]. UCI Machine Learning Repository, CC BY 4.0. Details on the data page.