Guide

Managing AI in a team that already works

The tools are not the hard part. Deciding who owns the output is.

8 min read·Back to guides
A small team working together at a table

Rolling AI into a working team fails in predictable ways, and none of them are technical. They are about ownership, measurement, and what happens the first time the output is wrong in front of a customer. Here is the management side, which is the side that decides whether any of it sticks.

1
Owner per output, always human
1
Shared library, not training
3
Failure modes, all in month two
2
Numbers worth tracking

One rule, and everything else follows

The person who sends the work owns the work. Not the tool, not the person who wrote the prompt, not the vendor.

This sounds obvious and it is routinely not said, which is why the first bad output produces a conversation about the model instead of a conversation about the review that did not happen.

Say it before anything is rolled out, and the rest of the policy stops being necessary. People do not send things they are accountable for without reading them. The teams that get into trouble are the ones where accountability quietly moved to the tool and nobody noticed until a customer did.

A shared library beats a training session

The instinct is to train everybody. What actually changes behaviour is a place where the prompts that work already live.

A prompt that reliably produces a good first draft of a support reply is an operating asset. It was written once, tested, corrected, and it now saves twenty minutes a day for everybody who has it. Left in one person's history, it saves twenty minutes for one person.

So keep a shared document. One entry per job: what it is for, the prompt, one example of good output, and who to ask. Ten good entries beat a two-hour workshop, because they arrive at the moment of need rather than three weeks before it.

And the person who finds a better version edits the entry. That is the whole governance model, and it is the same one that works for any other shared document.

Entry fieldWhy it is there
What job it doesSo somebody can find it without reading the prompt
The prompt itselfCopy and paste, no reconstruction
One good outputSo a reader can tell whether it worked before they trust it
Who to askSo an improvement has somewhere to go

The three failure modes, and they all arrive in month two

Week one goes well everywhere. The interesting failures are later.

Quiet quality drift. The output is fine, then acceptable, then slightly off, and nobody notices because each step is small and the reviewer is now skimming. The fix is a check on a schedule, not on a feeling: read five outputs properly, once a week.

The expertise gap. Junior people improve fastest with these tools, which is the point, and it also means they now produce work they could not have produced and cannot fully judge. That is a review question and a training question, not a tooling one.

Tool sprawl. Four people quietly subscribe to four different things, none of it is on the business tier, and your data rule now covers none of what is actually happening. This is the one that turns into an incident, and it is caused by the company account being slower to get than a personal one.

Two numbers worth tracking, and the ones that mislead

Time on the job, before and after, measured on a real sample rather than estimated. Ten instances timed by hand beats a survey.

The correction rate: what fraction of outputs needed a meaningful edit before they went out. If it climbs, quality is drifting. If it is near zero, either the job was easy or nobody is checking, and it is worth finding out which.

What misleads: usage. Calls per week, seats active, messages sent. Those numbers go up during any rollout and say nothing about whether the work got better or the job got faster. A team can be very busy with a tool that is costing it time.

Start with one job, out loud

Pick one job that one team does often, that costs real time, and where a wrong answer is embarrassing rather than dangerous. First drafts of replies. Summarising a call. Turning notes into a brief.

Run it for a month with one owner, the library entry written down, and the two numbers measured at the start and the end.

Then tell the whole company what happened, including if it did not work. The single biggest predictor of whether the second rollout goes well is whether people believe the first report, and the fastest way to lose that is to announce a success nobody in the team recognises.

Whoever's name is on the work still owns it. Every rule below is a consequence of that one, and a team that has not said it out loud will discover it the expensive way.

See the operations templates

Have a question?

We're here to help. Get in touch and let's talk.