Hallucination
A confident, plausible-sounding answer from an AI system that is factually wrong. Recognising one, and checking for one, is covered in every foundational course this site runs.

The word is unhelpful, because it suggests something visibly strange happening on screen. What actually happens is the opposite. The wrong answer arrives in the same register, at the same length and with the same confidence as a right one, which is exactly why it gets missed.
Why a wrong Answer sounds Right
A large language model is trained to predict and generate language, not to check it. Fluency is what it optimises for, so a fluent invention and a fluent fact are produced by the same process and leave no difference on the surface for a reader to notice. There is no hedge in the tone, no gap in the sentence, nothing that reads as a guess.
That single fact explains most of what follows on this page. A model has no internal sense of “I am not sure about this part.” It generates the next likely word given everything written so far, whether or not that word is attached to anything real, and a sentence built that way can be completely wrong while reading exactly like a sentence that is completely right. A model can produce a confident hallucination with exactly the same tone as a correct answer, and knowing that fact is what actually lets a reviewer catch a wrong result rather than approve it on confidence alone. It is a governance question, not a technical curiosity, because the fix is a checking habit rather than a smarter way of reading the tone.
The Specific, Checkable detail is the most Dangerous kind
Hallucination does not usually show up as an obviously strange or nonsensical sentence. It shows up as a specific, checkable-looking detail: a citation, a case number, a statistic, a quote attributed to a real person, a date. Each of these has a recognisable shape, and a model can reproduce that shape fluently with nothing real behind it. A properly formatted citation to a source that does not exist reads as more credible than a vague, hedged answer would, precisely because it looks sourced. That is the opposite of what most people expect a mistake to look like, and it is why a reader who is only watching for something that “sounds off” tends to miss it entirely.
This is also why the risk grows with the stakes of the task rather than shrinks. A generic first draft that goes wrong in a vague way gets caught on a read- through. A wrong statistic dropped into a client report, or an invented clause summarised from a contract that never actually said it, carries the same confident tone as everything correct around it, and it is the specific, plausible-looking detail that does the most damage precisely because nobody thought to question it.
The Knowledge cut-off makes it worse
A model is trained on material available up to a particular point in time. Anything that happened after that point is simply not in it, unless the tool has been separately connected to a live source, a technique called retrieval augmented generation. Ask an ungrounded model about a change made last week and one of two things happens. Either it says it does not know, which is the honest and useful answer, or it answers confidently from older material as though nothing has changed, which is a hallucination with a very specific cause: the gap between what actually happened and what the model was last shown.
That distinction matters for anyone deciding how to use a tool day to day. A model with no connection to current information is not being careless when it gets a recent event wrong. It is doing exactly what it was built to do, complete a plausible pattern, on the only material it has, which happens to be out of date. The fix is not a better prompt. It is giving the tool a live source to draw from in the first place.
Hallucination and Bias are not the same failure
The two get lumped together as “AI going wrong,” but they are separate problems with separate fixes. A hallucination is a fluent, confident statement that is simply invented, with nothing behind it. Bias is different: it is an answer that skews in a particular direction because the material a model was trained on skewed that way, on a subject, a group, a language or a viewpoint. A biased answer can contain no invented facts at all and still be a problem. A hallucinated answer can carry no bias at all and still be wrong. Both need a human check before the result goes anywhere, but it is not the same check. Catching a hallucination means asking whether a claim is actually true or actually in the source. Catching bias means asking whether the answer would look the same for a different group, a different name or a different starting assumption, which is a harder question to notice because a fluent, well- punctuated answer gives no more warning of one than it does of the other.
That distinction is why the ethics and hallucination material in the generative AI foundations sits alongside a separate conversation about bias, transparency and over-reliance, rather than folding both into one slide about “AI mistakes.” Treating them as the same problem tends to produce a review process that catches whichever one the reviewer happens to be thinking about that day, and misses the other.
Where the cost actually lands
The consequence of an unchecked hallucination is not evenly spread across a business. A generic internal draft that contains one is a wasted five minutes once somebody notices. A hiring recommendation, a client-facing summary of a contract, or a regulated communication that contains one is a different order of problem, because the invented detail carries the same authority as everything true around it and is read, acted on or sent out before anyone thinks to question it. This is the same reasoning behind naming a checkpoint for the highest-exposure use cases specifically, rather than applying one uniform level of review everywhere: the review has to be proportionate to what happens if the checked claim turns out to be invented.
Grounding narrows what there is to check
The structural answer is grounding, which ties an answer to a specific set of source material, usually the organisation’s own documents, rather than letting the tool answer from its training alone. Grounding does not remove the need to check. It changes what checking actually involves. Without grounding, a reviewer is asking an open-ended question: is this true, in general? With grounding, the reviewer is asking a much narrower one: is this actually in the source document, or is it not? The second question is one a person can usually answer in seconds. The first one is not, which is the entire reason grounding reduces risk even though it never eliminates it.
Checking is a habit, not a warning slide
A rule that says “a person should check the output” is not a control, it is a hope. A checkpoint that works names who checks it, at what stage, and against what standard, before the result reaches a client, a decision or a public output. That matters most where the exposure is highest, a hiring decision, a regulated communication, a client deliverable, and it has to be staffed and actually used under deadline pressure rather than exist only on paper. A checkpoint that only gets followed when there is time to spare is not a checkpoint, it is a preference.
The same oversight question gets harder once an AI agent is running unattended rather than answering one prompt at a time. A chatbot produces one answer for a person to read before acting on it. An agent can take several steps and act directly on a system with nobody reviewing each one, so an invented fact partway through a task does not get caught at the point it is produced. Oversight has to extend past chatbots to anything acting on an organisation’s systems without a person in the loop, which is exactly where the EU AI Act sets its own obligations to scale with the risk a particular use case actually carries.
Two practical habits do most of the remaining work, and both are teachable in an afternoon. Make the tool quote the source it drew an answer from, so a reviewer has something concrete to check against. And make it flag what it does not know rather than fill the gap with something plausible, so a gap in the material shows up as a gap on the page instead of as an invented answer that looks complete.
Where it is taught
Recognising a hallucination, and checking for one, is covered in every foundational course this site runs, not as a single warning slide but as a working habit built into the session. The ChatGPT programmes, the Copilot days, the leadership masterclass and the HR-specific session all deal with it directly, including the same two habits above: making a tool quote its source, and making it flag what is missing rather than guess. That is the same reason AI literacy on this site is built around the workflows staff already do rather than delivered as generic awareness content. A person who has checked a hallucinated answer once, on a task they recognise, carries that habit into every tool they use afterwards. A person who has only been warned about the concept in the abstract usually does not.
FAQ
Questions about Hallucination
Is an AI Hallucination the same thing as the Model lying?
No. Lying requires knowing the truth and saying something else on purpose. A large language model has no such knowledge to hide, only a prediction of what a plausible answer looks like. It produces a wrong statement the same way it produces a right one, by predicting the next likely word, so there is no intent behind either result. That is exactly why the tone gives nothing away.
Can Careful Prompting stop Hallucination happening?
It reduces it, not removes it. A clearer, more specific prompt narrows the range of plausible answers the model can generate, which lowers the chance of a wrong one, but it does not give the model a way to check a fact against a source. Grounding does that job. Prompting and grounding solve different halves of the same problem.
What is the difference between a Hallucination and Bias in an AI Answer?
They are separate failures. A hallucination is a fluent, confident statement that is simply invented. Bias is an answer that skews in a particular direction because the material the model was trained on skewed that way, on a subject, a group or a viewpoint. A biased answer can contain no invented facts at all, and a hallucinated answer can carry no bias at all. Both need a human check, but they are not the same check.
Does Grounding an AI Tool in our own Documents remove the Risk entirely?
No, and treating it as a full fix is itself a risk. Grounding ties an answer to a specific set of source material instead of the model's training alone, which narrows what there is to check. It does not remove the need to check. What it changes is the question a reviewer is answering: not is this true in general, but is this actually in the source document or not, which is a question a person can answer quickly and with confidence.
Why would an AI Tool invent something as Specific as a Citation or a case number?
Because the model is completing a pattern, not checking a record. A citation, a case number and a statistic all have a recognisable shape, and the model can reproduce that shape fluently without anything real behind it. The result reads as more credible than a vague answer would, precisely because it looks properly sourced, which is what makes this pattern more dangerous than an obviously wrong guess.
Who should be Responsible for catching a Hallucination before it reaches a client?
A named person, at a named stage, checking against a named standard, before the result goes anywhere. A general instruction to 'have someone check it' is not a control, it is a hope, and it tends to fail exactly when deadline pressure is highest. The checkpoint has to be staffed and actually used, not just written down, and it matters most for a hiring decision, a regulated communication or a client deliverable.
Is Hallucination only a Risk with Chatbots, or does it affect AI Agents too?
It affects both, and it matters more with an agent. A chatbot answers one prompt at a time with a person reading the result before acting on it. An AI agent can take several steps and act on a system without a person reviewing each one, so an invented fact partway through does not get caught at the point it is produced. That is why oversight has to extend past chatbots to anything acting on an organisation's systems unattended.
Is checking for Hallucination covered in your AI Training Courses?
Yes, in every foundational course this site runs, not as a single warning slide but as a working habit. The ChatGPT programmes, the Copilot days, the leadership masterclass and the HR-specific session all deal with it directly, including how to make a tool quote its source and flag what it does not know rather than fill the gap with something plausible.
What it means for a Business
Recognising one, and checking for one, is covered in every foundational course this site runs, so it is taught as a habit rather than as a warning at the end of a slide deck.
Ready to Put This to Work?
Tell us where your team is with AI and we will tell you honestly what would make the biggest difference.