Guide

Turn support tickets and reviews into decisions

Your customers already told you what to fix. It is sitting in the inbox, unread as a set.

8 min read·Back to guides
A support team reviewing conversations on a screen

Every company reads its tickets one at a time and almost none read them as a body of evidence. The value is in the second reading, and it is the one AI is genuinely good at: not answering the ticket, but telling you what four hundred of them have in common.

1
Pass a month
2
Labels per ticket, no more
1
Ranked list at the end
0
New tools to buy

Why the second reading is the valuable one

A ticket read on its own asks one question: how do I make this person whole. That is the right question at the time and it is why nobody sees the pattern.

Read four hundred at once and different questions become answerable. Which step do people get stuck on. Which promise on the site is being misread. Which feature is generating work rather than saving it. Which complaint arrived quietly for six months before somebody escalated.

This is a counting job on unstructured text, and counting unstructured text is precisely the thing that used to require a person for a week and now takes an afternoon.

Two labels per ticket, decided in advance

The failure mode is asking for a summary. You get four hundred summaries, which is the same pile in different words.

Ask instead for exactly two labels per ticket, from lists you write before you start.

The first is where in the journey it happened: signing up, first use, paying, a specific feature, cancelling. Five to eight options, no more.

The second is what kind of problem it is: it broke, it was confusing, it is missing, it was slow, the price was wrong, expectations were set wrong somewhere else.

Two axes, both closed lists, and a free-text field for a one-line quote. That gives you something you can count, cross-tabulate, and read back to the person who owns the fix.

What you ask forWhy closed lists
Where in the journeySo you can see whether the problem is one step or everywhere
What kind of problemSo broken and confusing do not end up in the same bucket
One quote, verbatimSo the count has a human sentence attached when you present it
Nothing elseEvery extra field halves the reliability of the two that matter

The monthly pass, start to finish

Export last month's tickets, reviews and cancellation reasons into one file. Strip names and emails: the labels do not need them, and this keeps the whole exercise on the safe side of your data rule.

Run them in batches with the two lists in the instruction and the quote field. Batches, not one at a time, and not all four hundred in a single call.

Spot-check thirty by hand. If more than three are labelled in a way you disagree with, your lists are wrong, not the model. Fix the lists and rerun.

Count. Sort by count. Take the top three to whoever owns them, with the quotes attached.

That is the whole exercise, and it should take an afternoon once the lists exist. It is worth putting in the calendar on the same day each month, because the value compounds when you can compare this month's counts to last month's.

What to do with the ranking

Not everything at the top of the list is a product fix, and treating it that way is how this exercise gets abandoned.

Things that broke go to engineering. Things that confused people go to whoever writes the interface copy, and they are usually the cheapest wins on the list. Things that are missing go on the roadmap with a count beside them, which is a far better argument than the loudest customer. Things where expectations were set wrong go to marketing, because the fix is on the page that made the promise, not in the product.

The last category is the one this exercise finds that nothing else does. A complaint that appears forty times and is nobody's fault is usually a sentence on a landing page.

Where it goes wrong

Open-ended labels. If the model may invent categories, you get four hundred categories and nothing to count.

No spot-check. Labelling is a judgement and judgements drift. Thirty by hand is cheap and it is the only thing standing between you and a confident, wrong ranking.

Running it once. One pass tells you what hurts today. Three passes tell you what is getting worse, and that is the number worth acting on.

And summarising instead of counting. If the output is prose, somebody has to read it and form an opinion, which is the work you were trying to remove.

The output is not a report. It is one sentence naming the thing you are fixing this month, and the count that justifies it.

See the operations templates

Have a question?

We're here to help. Get in touch and let's talk.