ops-agent: an AI agent that triages support tickets and waits for a person to approve
I built this to show what a careful operations agent looks like: one that follows a company's own rules, checks its own work, and never acts alone.
I built it in October 2026 with Java 21, Spring Boot 4.1, Spring AI 2.0, Postgres with pgvector, and React. It runs on Gemini, and the model can be swapped in config. Today it's a working demo with two fictional businesses.
The problem
A support team gets the same kinds of tickets every day: refunds, broken items, late parcels, double charges, repairs. Each one needs the same steps. Read it, work out what it is and how urgent, check the policy, look up the order, decide who should handle it, and write the reply.
AI can do most of that quickly. The hard part is trusting it. A model will happily promise a refund that a manager still has to approve, or miss that a "colour swap" request mentions a plug that sparked. So I didn't want to build an agent that is right most of the time. I wanted one where the mistakes are caught before a customer sees them.
How it works
It reads incoming operations tickets and sorts them by type and urgency. When it needs more context, it looks up data in your own systems through tools, for example an order database or an internal docs page. Those lookups are read-only. Then it drafts the next action, such as a reply or a refund request, and lists the facts it relied on with their sources.
Then code checks all of that, whatever the model said. The result is saved as waiting for approval. A reviewer opens it, sees the draft next to the facts behind it, and approves, edits or rejects. Only an approved reply goes to the outbox.
A team can also use the agent from their own AI assistant, such as Claude, through an MCP server. The assistant can triage tickets and read results, but there is deliberately no approve tool. An assistant can be talked into things, by its user or by text inside a ticket, so approving stays with a person looking at the evidence.
Guardrails enforced in code
The prompt tells the model the rules. But some rules are too important to leave to the model, so code checks them after it answers:
- Refund limits. A support agent can approve refunds up to $100 on its own. Anything above goes to whoever can approve it, and the code adds that escalation if the model forgot.
- Refunds checked against real data. For every proposed refund, the code fetches the order and payments itself and checks the amount isn't more than what's left to refund, and that shipping is only refunded when the policy allows it.
- No promises that still need approval. If the draft says "I'm issuing your refund now" while a manager still has to approve it, the reviewer gets a warning with better wording.
- Urgent tickets always go to someone. If the model marks a ticket urgent but hands it to nobody, the code sends it to the business's on-call role.
- No internal details in drafts. Role names, priority codes, error codes and internal limits are swapped for customer wording or flagged.
- Every fact has a source. A fact or source with nothing behind it is flagged, and only documents that were actually retrieved can be cited.
Each check gives the reviewer a typed warning, and the warnings set the risk level they see first. High risk means a policy flag or an urgent ticket. The reviewer still decides.
Hi Rosa,
I can certainly help you set up a return for your faulty kettle and issue a full refund, including your express shipping fee.
Thanks,
Larkfield Goods support
Caught in code. Ticket LH-005 asked for $110.95 back for a faulty kettle, just over the $100 limit. The model would have handled it alone; code sent it to the team lead and flagged the missing payments lookup.
Not caught by a check. The draft still promises a full refund the team lead hasn't approved. That's why a person approves every reply.
How I measure it
A demo shows the agent on the tickets you choose. An evaluation shows it on tickets you didn't.
I wrote sample tickets for each business with the answer I expected for each one: category, priority, who it goes to, whether it's in scope, and for returns, the reason. Every expected answer cites the policy lines behind it. The eval runs the agent on every ticket and scores each of those decisions.
The first 28 samples were used to tune the prompt and the rules, so the agent scores 100% on them. That says more about the tuning than about the agent, so I don't quote it.
For an honest number I wrote a separate holdout set: 15 new tickets, messier on purpose, with typos, two problems in one message, a wrong order number, a forwarded email chain, and a refund just over a limit. I wrote the expected answers from the policies alone, before any model run. Holdout tickets are never used for tuning. If one leads to a fix, it's marked as used and drops out of the score for good.
I never change an expected answer to raise a score.
The result
94% of decision fields right on 15 unseen tickets: 66 of 70 fields, and 13 of 15 tickets fully right.
| Business | Fields right | Tickets fully right |
|---|---|---|
| Online store | 46/50 | 8/10 |
| Property management | 20/20 | 5/5 |
| Both | 66/70 | 13/15 |
First run, 10 October 2026, on gemini-3.5-flash-lite. This is the published score and it isn't updated.
The two tickets it got wrong:
- A café asking to buy 24 mugs was treated as a sales pitch, because the store's wording only counted people asking about "their order" as customers.
- A request to swap a milk frother's colour mentioned that it smelled of burning plastic and the plug sparked. The agent rated it urgent but didn't hand it to anyone.
Both are fixed now. The wording counts anyone buying from the store as a customer, and urgent tickets always go to someone, in code. Fixing them used up those two tickets, so 94% stays the published score and isn't updated.
And the most important line: every one of those drafts still waits for a person to approve it.
Two businesses, no code changes
The agent runs two made-up businesses from the same code:
- Larkfield Goods, a small online store, with an order system, payments, and a refund policy with approval limits.
- Harbourline Property Management, which manages rented apartments. It has no order system, so the agent works from its documents only.
I'm keeping it configurable, so the same agent can work with any business's docs and systems. Each business is one profile file: its knowledge base, its categories, who handles what, its limits and its tools. You pick the profile, and the agent works with those. The app checks the profile when it starts and won't run until every problem is fixed. The code and the prompt have no business-specific values of their own.
Adding Harbourline took a profile, a folder of policies and sample tickets. Its eval scored 32/32 on the samples and 20/20 on the holdout.
The reviewer desk
A reviewer's job is to decide, in under a minute, whether a draft can go out. So everything on the screen serves that decision. The risk level shows its first reason at the top. Approve and reject are always at the bottom with a key each. Approving waits six seconds before anything happens, so Undo really undoes it.
Security
The demo is built to be deployed, so it's protected like a real app. Reviewers log in, passwords are stored only as hashes, and every change needs a CSRF token. Repeated wrong passwords lock out that username from that address. The admin side and the AI assistant each have their own key. Every model call counts toward a daily limit kept in the database, so a looping assistant can't run up a bill. Ticket text is always shown as plain text, and the page runs no scripts but its own. Before publishing, I scanned the whole git history for secrets and found none.
What I'd do differently
- Write the holdout set on day one. I wrote it after the tuning, and until then my only score was 100% on tickets I'd tuned against, which told me very little.
- Have the model say which fact each sentence uses. Today the desk links sentences to facts by matching shared terms, like an order number or an amount. It never draws a false link, but it misses some real ones. The model stating the link would be better.
- Check claims, not just numbers. The checks catch an amount that isn't in the order, but not a sentence that changes a fact's meaning while keeping its numbers. In the gas replay, the policy's "the engineer attends within 2 hours" became "you will hear back within 2 hours", and nothing flagged it. That needs a proper claim check.
- Design the review screen around the decision first. My first version looked like a dashboard. People told me it looked auto-generated, and they were right. The second version starts from what a reviewer does in that one minute.
How it was built
I built ops-agent quickly with AI-assisted development. The architecture, the rules, the review of every change, and the evaluation are mine. The code is private; I share read access with serious prospects.