Skip to content
Uzair Azhar
Menu

Customer Support Agent

A tool-calling agent that answers order questions from the retailer's real customers about their real invoices. It looks things up, applies the refund policy in code, cites the policy, and stops for a person before any refund over £50. Every step is shown.

LiveAI and LLM appsReal invoicesSimulated: shipments, policy, messages

1Problem

Most customer messages to an online shop are about one order: something arrived broken, something is missing, a parcel is late. Answering them means opening the invoice, checking tracking, reading the policy and working out what is owed. It is repetitive, and mistakes cost money in both directions.

Chatbots that answer from a prompt alone make it worse: they promise refunds the policy does not allow, invent order details, and can be talked into things.

2Why it matters

An agent is only useful in operations if it is safe to let it act. That means its facts come from systems, its amounts from rules, its actions are limited and reviewable, and a person decides anything significant. Those are the properties this project demonstrates, with a trace that lets anyone check what happened.

3Solution

The agent works through tools. It reads the invoice and the tracking, searches the policy, and asks the refund rules what is due. It can record a refund only for the exact amount the rules returned: up to £50 it does so itself; above that, it asks for approval and the conversation waits for the supervisor. It opens tickets for work a person must do and drafts a reply that cites the policy. Nothing is sent and no money moves.

A free-tier model drives the loop when one is configured. Without one, a scripted agent takes the same steps with the same tools, so the demo always works and the difference between the two can be measured.

4Architecture

Customer message

About one of their real invoices

Agent loop

Up to 8 steps and 25 seconds a turn. Each step: the model sees the conversation and the tools, proposes calls; arguments are validated, tools run, results go back.

Provider chain

Groq, then Cerebras (free tiers, when keys are set), then the scripted agent. Token budgets per day and per visitor.

  • get_invoice

    Real invoice lines and totals

  • list_customer_invoices

    The customer's recent orders

  • get_shipment_status

    Simulated tracking for the invoice

  • search_policy

    Full-text search, numbered passages

  • calculate_refund

    Policy rules in code: eligibility and amount

  • request_approval

    Records a refund; over £50 waits for a supervisor

  • create_ticket

    Hands work to a person

  • draft_reply

    Saves the reply; nothing is sent

Real. Customers (anonymous IDs), invoices, products, quantities and prices from UCI Online Retail II.

Simulated. Shipments, the service policy, tickets, refunds and the messages themselves. Nothing leaves the system.

Human in the loop. A refund above £50 pauses the conversation until the supervisor (you) approves or rejects it.

A customer message goes to an agent loop that calls a provider chain: free-tier models when configured, then a scripted agent. The agent uses eight tools over real invoices, simulated shipments and a written policy; refunds above £50 wait for a supervisor.

5Technical implementation

Agent loop
Written by hand against the OpenAI-compatible chat API; no agent framework
Models
Groq and Cerebras free tiers when configured, a scripted agent as the last step
Tools
Eight Python functions with Pydantic argument models, scoped to one customer
Data
Real invoices in a read-only DuckDB snapshot; policy search in PostgreSQL
API
FastAPI with per-visitor quotas; conversations kept for 24 hours
Interface
Next.js: chat, supervisor approval and a step-by-step trace

The loop. Each step sends the conversation and the tool definitions to the provider chain. All tool calls in a response run as a batch (looking up the invoice and its tracking happens in one step), each result is returned to the model, and the loop ends when the agent drafts a reply or a refund needs approval. A supervisor's decision is added to the conversation and the loop resumes.

Refund rules. Damaged or missing items within 14 days of delivery, late deliveries (more than 3 days after the estimate: the shipping charge, or 5% of the goods up to £10), lost parcels (10 days past the estimate) and returns within 30 days (refunded on receipt). Credit notes are never refunded again.

Real and simulated data. Invoices come from 797,885 real invoice lines of customers with an ID, after removing the source export's duplicate lines. Shipments are simulated from each invoice's number and date, so the same invoice always has the same tracking; carrier services have generic names. The desk works as of 10 December 2011, the day after the data ends.

Policy search. The policy is split into passages by section and searched with PostgreSQL full-text search through the shared retrieval package; passages are numbered so replies can cite them, and a reply cannot cite a passage the agent did not retrieve.

6Live demonstration

Pick a scenario, or write your own message as one of the customers. The left side shows the conversation; the right side shows every step the agent took. When a refund needs approval, you decide.

Loading scenarios…

Portfolio demonstration. Customers and invoices are real (anonymised by the source); shipments, the policy, tickets and messages are simulated. Each visitor can start 20 conversations an hour.

7Results

Results are read from the running system and are not available right now. They appear after the first scheduled run of the scenario suite.

8Failure handling

The model proposes invalid arguments or an unknown tool
Arguments are validated before anything runs. The error goes back to the model as the tool result, so it can correct itself on the next step.
The model, or the customer, tries to raise a refund
Refund amounts come only from calculate_refund. request_approval refuses any amount that differs from the rules' result by more than a penny, and refuses to run before a calculation.
A customer asks about someone else's order
Every tool is scoped to the conversation's customer. Another customer's invoice gets the same answer as an invoice that does not exist, so nothing leaks.
The agent loops or stalls
A turn stops after 8 steps or 25 seconds. The conversation is then handed to a person with a ticket, and the customer is told so.
No model is available, or the free tier is used up
The provider chain falls through to the scripted agent, which uses the same tools and rules. The trace says which provider answered each step and why earlier ones were skipped.
A large refund needs a decision
Above £50 the agent can only request approval. The conversation pauses, the visitor decides as supervisor, and the agent drafts the follow-up from the decision.
Anything would reach a real person or account
Nothing can: refunds, tickets and replies are records in this demo's database, deleted after a day.

9Deployment

The agent runs in the site's API container. The order snapshot is built into the image from the verified dataset download; the policy is indexed into PostgreSQL the first time it is searched. Conversations are stored for 24 hours and removed by a daily job, which also re-runs the scripted baseline.

nginx gives the agent's routes a longer read timeout (a turn with a model can take several seconds) and a stricter request rate than the rest of the API.

10Cost

$0 in API spend. Only free tiers are ever used, through a token ledger with daily limits per provider and per visitor. When the limits are reached the chain falls through to the scripted agent, which costs nothing.

Every model step resends the instructions and the eight tool definitions, about 5,600 characters (roughly 1,400 tokens by the ledger's estimate of four characters per token), plus the conversation and tool results so far. The trace shows the measured tokens of each step when a model answers.

11Limitations

  • The scripted agent reads keywords and product names. It handles the scenarios it was written for and simple variations; anything else gets a clarifying question or a hand-over.
  • Models have not been scored on this server yet, because none is configured by default. The suite is ready to run when a key is.
  • Shipments and the policy are simulated, and credit notes in the source data do not reference the invoice they credit, so the agent cannot check for an earlier refund on the same items.
  • Each conversation is one customer; there is no login. The customer is chosen in the demo, which a real system would take from an authenticated session.

12Source code

The full source is on GitHub under the MIT licence, with the tests and scenarios.

Data: Chen, D. (2012). Online Retail II [Dataset]. UCI Machine Learning Repository, CC BY 4.0. Details on the data page.