Skip to content
Rajkumar Vallepu

ops-agent: an AI agent that triages support tickets and waits for a person to approve

I built this to show what a careful operations agent looks like: one that follows a company's own rules, checks its own work, and never acts alone.

I built it in October 2026 with Java 21, Spring Boot 4.1, Spring AI 2.0, Postgres with pgvector, and React. It runs on Gemini, and the model can be swapped in config. Today it's a working demo with two fictional businesses.

The problem

A support team gets the same kinds of tickets every day: refunds, broken items, late parcels, double charges, repairs. Each one needs the same steps. Read it, work out what it is and how urgent, check the policy, look up the order, decide who should handle it, and write the reply.

AI can do most of that quickly. The hard part is trusting it. A model will happily promise a refund that a manager still has to approve, or miss that a "colour swap" request mentions a plug that sparked. So I didn't want to build an agent that is right most of the time. I wanted one where the mistakes are caught before a customer sees them.

How it works

01 ticket inreviewer screensor an ai assistant 03 triage agentsearches docs,looks up, drafts 04 guardrailscode checks thedraft, sets risk 05 approval gatea person approvesor rejects 06 outboxapproved repliesonly 02 businessprofilecategories,roles, limits,docs, tools no approve tool for ai assistants; order lookups are read-only ops-agent overview · rev 01 drawn: RV · 2026-10
A ticket comes in from the reviewer screens or a team's AI assistant. The business profile sets up the agent and the checks. The triage agent searches the docs, looks up orders read-only and drafts a reply. Code checks the draft and sets a risk level. A person approves or rejects at the gate, and only an approved reply reaches the outbox.

It reads incoming operations tickets and sorts them by type and urgency. When it needs more context, it looks up data in your own systems through tools, for example an order database or an internal docs page. Those lookups are read-only. Then it drafts the next action, such as a reply or a refund request, and lists the facts it relied on with their sources.

Then code checks all of that, whatever the model said. The result is saved as waiting for approval. A reviewer opens it, sees the draft next to the facts behind it, and approves, edits or rejects. Only an approved reply goes to the outbox.

A team can also use the agent from their own AI assistant, such as Claude, through an MCP server. The assistant can triage tickets and read results, but there is deliberately no approve tool. An assistant can be talked into things, by its user or by text inside a ticket, so approving stays with a person looking at the evidence.

Guardrails enforced in code

The prompt tells the model the rules. But some rules are too important to leave to the model, so code checks them after it answers:

  • Refund limits. A support agent can approve refunds up to $100 on its own. Anything above goes to whoever can approve it, and the code adds that escalation if the model forgot.
  • Refunds checked against real data. For every proposed refund, the code fetches the order and payments itself and checks the amount isn't more than what's left to refund, and that shipping is only refunded when the policy allows it.
  • No promises that still need approval. If the draft says "I'm issuing your refund now" while a manager still has to approve it, the reviewer gets a warning with better wording.
  • Urgent tickets always go to someone. If the model marks a ticket urgent but hands it to nobody, the code sends it to the business's on-call role.
  • No internal details in drafts. Role names, priority codes, error codes and internal limits are swapped for customer wording or flagged.
  • Every fact has a source. A fact or source with nothing behind it is flagged, and only documents that were actually retrieved can be cited.

Each check gives the reviewer a typed warning, and the warnings set the risk level they see first. High risk means a policy flag or an urgent ticket. The reviewer still decides.

Hi Rosa,

I can certainly help you set up a return for your faulty kettle and issue a full refund, including your express shipping fee.

Thanks,
Larkfield Goods support

Caught in code. Ticket LH-005 asked for $110.95 back for a faulty kettle, just over the $100 limit. The model would have handled it alone; code sent it to the team lead and flagged the missing payments lookup.

Not caught by a check. The draft still promises a full refund the team lead hasn't approved. That's why a person approves every reply.

How I measure it

A demo shows the agent on the tickets you choose. An evaluation shows it on tickets you didn't.

I wrote sample tickets for each business with the answer I expected for each one: category, priority, who it goes to, whether it's in scope, and for returns, the reason. Every expected answer cites the policy lines behind it. The eval runs the agent on every ticket and scores each of those decisions.

The first 28 samples were used to tune the prompt and the rules, so the agent scores 100% on them. That says more about the tuning than about the agent, so I don't quote it.

For an honest number I wrote a separate holdout set: 15 new tickets, messier on purpose, with typos, two problems in one message, a wrong order number, a forwarded email chain, and a refund just over a limit. I wrote the expected answers from the policies alone, before any model run. Holdout tickets are never used for tuning. If one leads to a fix, it's marked as used and drops out of the score for good.

I never change an expected answer to raise a score.

The result

94% of decision fields right on 15 unseen tickets: 66 of 70 fields, and 13 of 15 tickets fully right.

Holdout result, first run, 10 October 2026, gemini-3.5-flash-lite
BusinessFields rightTickets fully right
Online store46/508/10
Property management20/205/5
Both66/7013/15

First run, 10 October 2026, on gemini-3.5-flash-lite. This is the published score and it isn't updated.

The two tickets it got wrong:

  • A café asking to buy 24 mugs was treated as a sales pitch, because the store's wording only counted people asking about "their order" as customers.
  • A request to swap a milk frother's colour mentioned that it smelled of burning plastic and the plug sparked. The agent rated it urgent but didn't hand it to anyone.

Both are fixed now. The wording counts anyone buying from the store as a customer, and urgent tickets always go to someone, in code. Fixing them used up those two tickets, so 94% stays the published score and isn't updated.

And the most important line: every one of those drafts still waits for a person to approve it.

Two businesses, no code changes

The agent runs two made-up businesses from the same code:

  • Larkfield Goods, a small online store, with an order system, payments, and a refund policy with approval limits.
  • Harbourline Property Management, which manages rented apartments. It has no order system, so the agent works from its documents only.

I'm keeping it configurable, so the same agent can work with any business's docs and systems. Each business is one profile file: its knowledge base, its categories, who handles what, its limits and its tools. You pick the profile, and the agent works with those. The app checks the profile when it starts and won't run until every problem is fixed. The code and the prompt have no business-specific values of their own.

Adding Harbourline took a profile, a folder of policies and sample tickets. Its eval scored 32/32 on the samples and 20/20 on the holdout.

The reviewer desk

A reviewer's job is to decide, in under a minute, whether a draft can go out. So everything on the screen serves that decision. The risk level shows its first reason at the top. Approve and reject are always at the bottom with a key each. Approving waits six seconds before anything happens, so Undo really undoes it.

The review screen: the queue of waiting tickets on the left, and the open ticket on the right with the customer's message, the draft reply and the facts it rests on.
The queue and the open ticket on one screen. The draft is the largest text, and the facts it rests on sit beside it.
The same review screen in the dark theme.
Light and dark, and every action has a key.
Close-up of a draft sentence being pointed at: a thread links it to the policy fact it rests on, with the fact's source underneath.
Point at a sentence to see the fact behind it and where that fact came from. A sentence with no link is the one to read closely.
Close-up of the Before you approve list: a refund over the agent's limit marked as a policy flag, then four warnings, and what approving will do.
The risk level and its first reason sit at the top. The rest are listed under "Before you approve".

Security

The demo is built to be deployed, so it's protected like a real app. Reviewers log in, passwords are stored only as hashes, and every change needs a CSRF token. Repeated wrong passwords lock out that username from that address. The admin side and the AI assistant each have their own key. Every model call counts toward a daily limit kept in the database, so a looping assistant can't run up a bill. Ticket text is always shown as plain text, and the page runs no scripts but its own. Before publishing, I scanned the whole git history for secrets and found none.

What I'd do differently

  • Write the holdout set on day one. I wrote it after the tuning, and until then my only score was 100% on tickets I'd tuned against, which told me very little.
  • Have the model say which fact each sentence uses. Today the desk links sentences to facts by matching shared terms, like an order number or an amount. It never draws a false link, but it misses some real ones. The model stating the link would be better.
  • Check claims, not just numbers. The checks catch an amount that isn't in the order, but not a sentence that changes a fact's meaning while keeping its numbers. In the gas replay, the policy's "the engineer attends within 2 hours" became "you will hear back within 2 hours", and nothing flagged it. That needs a proper claim check.
  • Design the review screen around the decision first. My first version looked like a dashboard. People told me it looked auto-generated, and they were right. The second version starts from what a reviewer does in that one minute.

How it was built

I built ops-agent quickly with AI-assisted development. The architecture, the rules, the review of every change, and the evaluation are mine. The code is private; I share read access with serious prospects.

Play back three real tickets

Contact

If your team answers the same questions all day, I'd be glad to show you how this would work with your own policies. Please do reach out.

rajkumar.vallepu1997@gmail.com