How to check AI output before it goes to a customer
A practical way to review AI drafts before customers see them: what to check, in what order, how much checking each task needs, and who does it.

AI text should never reach a customer on the strength of how it sounds. It reads fluently when it is right and just as fluently when it is wrong, so the only reliable test is a check against something the AI did not write: the source document, the account record, the policy.
This article sets out a way to do that check. It covers where in a process the check belongs, what a reviewer should look at and in what order, how to decide how much checking each task deserves, and how a small team can run it without a quality department.
Key takeaways
- Put the check before the output does anything: before a message is sent, a record is changed or a refund is decided.
- Match the depth of review to the cost of being wrong, not to how good the draft looks.
- Check claims against the original source, never against the reviewer’s memory or the AI’s own summary.
- Use a short, fixed checklist in a fixed order: facts, numbers and names, commitments, tone, exposure of sensitive information.
- Name the reviewer. A review step with no owner quietly stops happening.
- Start strict, record what reviewers change, and loosen the checks only when the record supports it.
Why fluent output is the problem
Most people check AI output the way they would check a colleague’s draft: they read it and see whether it makes sense. That works on a human draft because human errors tend to show up as clumsy writing. An error in AI output does not. A hallucination arrives in the same polished sentences, with the same confident tone, as a correct statement. One published guide to fact-checking AI content lists the red flags as false information that sounds true, outdated or fabricated data, contradictions within a single piece, and a lack of context or sensitivity. Only some of those are visible on a read-through.
So the first change is to stop reading for sense and start reading for claims. Every sentence in a customer message that states something checkable (a price, a date, a delivery window, a policy, a name, a figure) is a claim, and each claim is unverified until a person has traced it back to its source. The same guide recommends doing this literally: look for the citation, search for the exact phrase in the source, and cross-check against a trusted reference. If the AI offers a link, open it. A link can be wrong, or can exist and say something different from what the text claims.
The second change is to accept that the AI cannot tell you which of its answers are wrong. Asking it to double-check itself is better than nothing, but it is not a control, because the same system that made the mistake is marking its own work. The control has to sit outside the model: a person, or a rule the model does not get to write.
This is also why the design of the workflow matters before any review begins. A workflow that answers only from approved material, a design known as grounding, gives the reviewer something definite to compare against. A workflow that answers from the model’s general knowledge gives the reviewer nothing to compare against except their own recall, which is the weakest standard available.
Where the check belongs
A check is only useful if it happens before the harm. A published guide on review points puts it plainly: place the checkpoint before a message is sent, a record is changed, or a case is approved, denied, refunded or routed. It is a simple rule, and it removes an argument that otherwise recurs, which is whether a quick look afterwards is good enough. After a customer has read the message, you are no longer reviewing it. You are correcting it.
In practice this means finding the moment in each process where the AI’s output stops being a draft and becomes an action. Three moments cover most small businesses.
Before a customer sees a reply. A reply drafted by AI, whether it is an email, a chat answer or a quote cover note, is held until a named person releases it. This is the most visible checkpoint and the one most worth keeping strict.
Before a label drives a decision. When AI sorts enquiries, tags complaints or scores leads, the label decides who sees the item and how fast. A wrong label rarely looks like an error. It looks like a customer who waited a week. Check a sample of labels, and check them on the days that matter most.
Before a summary is acted on. When a person makes a decision from an AI summary of a contract, a thread or a report, the summary has become the evidence. Spot-check summaries against the original, especially where a figure or an obligation is mentioned.
The pattern repeats across the guidance on customer service in particular. Most mature set-ups use some mix of review before the response goes out, correction after it, and escalation to a person when the AI is unsure. For a small business, the first and third are the ones worth building. Correction after the fact is what you do when the first two failed. Keeping a human in the loop at the right moment is the whole design.
What a reviewer should check, in order
A checklist works when it is short enough to use every time. The guidance on review points says most teams need only a few questions, and that is right: a long list is skipped under pressure, and a skipped check is worse than a short one done properly. The order below is deliberate. It moves from the most damaging error to the least, so that a reviewer who is interrupted has still done the part that matters most.
1. Do the facts match the source? Open the document the answer is supposed to come from and compare. Not the AI’s summary of it, the thing itself. If the draft states a policy, find the policy. If it quotes a status, look at the record.
2. Are the numbers, names and dates right? These are the details an AI is most likely to alter slightly, and the ones a customer is most likely to act on. Read each one against the source individually. A transposed date or a wrong name is easy to miss in a paragraph and expensive once sent.
3. Does it promise anything? Look for commitments the business has not made: a refund, a deadline, an exception, a guarantee. The guide on review points lists unsafe promises explicitly. AI is agreeable by design, and that tendency can produce generosity the business never authorised. A sentence such as “we will make sure this is resolved by Friday” is a commitment, and someone has to own it.
4. Is anything missing? A draft can be accurate in every sentence and still leave out the one thing the customer needed to know, such as the condition attached to an offer or the next step they have to take. Reread it as the customer, with only this message in front of you.
5. Does the tone suit this person? The same reply that suits a routine query can land badly with a customer who is already frustrated. Tone is the check where a human is hardest to replace, because it depends on knowing what has already happened between the business and this customer.
6. Does it expose anything it should not? Check for personal data, internal notes, other customers’ details or anything copied from an internal document. This is the check that can turn a wording problem into a data-protection one, so it belongs on the list even when it fails rarely.
If a draft fails at any point, do not patch the sentence and send it. Note what failed. A repeated failure at step one means the AI is answering from the wrong material. A repeated failure at step three means its instructions allow it to promise things. Both are workflow fixes, and the review step is where you find out that they are needed.
How much checking a task deserves
Not everything needs the full list. A business that reviews every output at full depth will either run out of reviewer time or stop reviewing. The question that sorts the work is the one from the guidance on review points: what happens if this output is wrong?
Where one bad output could cause real harm, the review is complete and every item is read before it goes out. That covers anything touching money, a contract, a legal obligation, health or safety, or a person’s data. This is the same line that separates automating a task from automating a decision: the more a step commits the business to something, the more a person has to stand behind it.
Where the output is low-risk, such as an internal brainstorm, a first-draft outline or a note that only the team will see, a spot check is enough. One published guide suggests sampling around 10 percent of items each day for lower-risk work, and testing with 10 to 20 real examples before going live. Those are a starting point to adjust, not a rule, but they show the right scale: a small, regular sample, taken from real work and not from tidy test cases.
Some situations should send the item to a person regardless of how it was classified. The customer-service guidance names the common ones: the AI is not confident in its answer, the customer is frustrated or upset, the query involves a refund, an account change or a medical or legal matter, the same question keeps coming back, or the customer asks for a human. Build these in as automatic hand-offs. A business that relies on a reviewer noticing them will miss the ones that arrive on a busy afternoon.
One principle keeps all of this honest: start strict and relax later. A review that is tightened after a bad week is a reaction. A review that begins strict and is loosened because the record shows it catching nothing is a decision. The second is much easier to defend.
Who does the checking, and how to know it works
In a small business there is rarely a quality team, so the practical question is who. The right reviewer is the person who knows the subject best and who is answerable for what the customer is told. That is usually whoever owns the process, with a named back-up for when they are away. It is not whoever happens to have spare time. This is the same lesson as in what happens when nobody owns the AI that has been rolled out: a step with no named person attached to it is a step that stops happening without anyone deciding to stop it.
Give the reviewer what they need in one place: the AI’s draft, the source it was meant to use, and the original customer message. A reviewer who has to hunt for the source will skip it. Ask them to record what they changed and why, even in a line. That record is the evidence that the step is worth its cost, and it is also how the workflow improves.
Review the record every few weeks. The measures worth watching are simple: how often reviewers change the draft, which kinds of error repeat, how long a review takes, and whether two reviewers disagree on the same item. If reviewers almost never change anything for a long stretch, the step may be loosened for that task. If the same mistake keeps returning, more checking is the wrong answer. Fix the instructions, the source material or the design of the workflow so that the mistake stops being produced.
The cost of all this is real and belongs in the budget. Reviewer time is one of the largest lines in the running cost of an AI workflow, and it is the one most often left out of the estimate. A review step that is so slow that it cancels the saving defeats the purpose, which is why a well-scoped workflow, with a narrow job and a clear point where a person steps in, is also the cheapest to check.
Finally, the habit has to be written down somewhere the business can point to. A short statement of who reviews what, and what must never be sent without a person reading it, is the practical core of AI governance for a small team. It does not need to be long. It needs to exist, and to be followed.
What to do next
Pick one place where AI output already reaches a customer, or soon will. Write down what happens if it is wrong, name the person who checks it, and give them the six questions above and the source to check against. Run it strictly for a month and keep the record of what they change. If you want a second opinion on where review belongs in a specific process, a scoping conversation is a quick way to find out before anything goes live.
Sources
- Human review points in AI workflows: where to check: sets out placing review before a message is sent or a case is decided, a risk-based approach, a short reviewer checklist, and the suggestions to test with 10 to 20 real examples and sample 10 percent of lower-risk items daily.
- Human-in-the-loop patterns for AI customer service in production: describes review before and after a response and escalation to a person on low confidence, negative sentiment or high-stakes actions.
- How to fact-check AI content like a pro: lists five fact-checking steps (citations, trusted sources, inconsistencies, timeliness, experts) and the red flags that signal unreliable AI content.
FAQ
Questions we get asked
How do you check AI output before sending it to a customer?
Check it in a fixed order: facts against the original source, then numbers, names and dates, then any promise or commitment the text makes, then tone, then whether it exposes anything it should not. Match the depth of checking to the cost of being wrong. A message that touches money, a contract or personal data is read in full by a person before it goes out.
Does every piece of AI output need Human review?
No. Review should follow risk. Internal drafts and low-stakes labelling can be spot-checked, while anything customer-facing, financial, legal or involving personal data should be reviewed before it is sent. A published guide on review points suggests starting strict and relaxing the checks only once results have stayed stable for a while.
What is the most Common mistake in checking AI output?
Reading it for how good it sounds instead of whether it is true. AI text is fluent even when it is wrong, so a reviewer who only reads for sense will miss an invented figure or a wrong date. The fix is to check claims against the source document, not against the reviewer's memory.
Who should review AI output in a Small Business?
The person who knows the subject best and who is accountable for what the customer is told, not whoever has spare time. For a small team that is usually the owner of the process, with a named back-up. Naming a reviewer matters more than having a formal quality team.
How do you know the review step is working?
Track how often reviewers change the AI's draft, which kinds of error repeat, and how long each review takes. If reviewers almost never change anything, the step may be loosened. If the same error keeps returning, the fix belongs in the workflow itself and not in more checking.
Ready to put this to work?
Tell us where your team is with AI and we will tell you honestly what would make the biggest difference.

