Large Language Model (LLM)
The technology behind tools such as ChatGPT and Copilot: a model trained on a very large amount of text to predict and generate language. It is the basis both for the courses we teach and for the workflows an audit identifies.

What a Large Language Model actually is
A large language model is a system trained on a very large amount of existing text, pulled from books, articles, websites and other written material, to learn the statistical relationships between words: which ones tend to follow which others, in which order, in which context. It is not a database of stored answers. Nothing it produces is looked up. It is generated, one piece at a time, based on patterns learned during training.
“Large” refers to the scale of both the training material and the model itself, measured in the number of adjustable values, called parameters, it holds. That scale is what lets it hold enough of the structure of ordinary language to produce a fluent, well-formed sentence about almost any subject a person might raise with it, not just the narrow set of topics it was deliberately built to handle.
The name is precise about one more thing. A large language model is built on language specifically. An image generator learns patterns in pictures instead of words and is not a language model, even though it sits in the same broader category. That category is generative AI, and it is the umbrella term for any system that produces new content rather than only sorting or retrieving what already exists.
How it produces an Answer
Training happens once, in advance, and is not repeated for each conversation. During that process the model reads enormous quantities of text and adjusts its internal parameters so that, given the start of a sentence, it gets progressively better at predicting a plausible next piece.
Answering a prompt uses the same underlying skill. The model reads the prompt, then predicts the most likely next token, the small chunk of text roughly the size of a word or a piece of one, adds it to what has been written so far, and predicts the next one after that. It repeats this step by step until a full answer has been built. It is not searching for the answer. It is completing a pattern, one token at a time.
The amount of text a model can hold in view while doing this, the prompt, an attached document and its own growing answer together, is called its context window. This is a real, published figure for some tools rather than an abstract idea: Google states one million tokens for Gemini, extended directly into tasks such as querying a large PDF stored in Drive without splitting it into smaller pieces first. A larger context window means a model can work across a longer document, or a longer back-and-forth conversation, without losing track of what was said earlier in it. Microsoft, by contrast, publishes no equivalent token figure for Copilot, tracking usage through AI credits and per-app limits instead, which is a difference worth knowing when comparing tools rather than assuming they measure capacity the same way.
Where the Technology actually shows up
Most of the named AI tools a business is likely to already use are built on a large language model at their core. ChatGPT, Claude, Gemini and Microsoft 365 Copilot are all products wrapped around one, with a chat interface, memory of the current conversation and a set of features each vendor chose to ship layered on top.
That wrapping is where the products actually differ, since the underlying skill, predicting the next plausible piece of text, is shared by all of them. Copilot’s distinguishing feature is grounding in a user’s own emails, chats, meetings and documents through Microsoft Graph, what Microsoft describes as Work IQ, rather than one open document at a time. Gemini’s is native presence across Google Workspace’s own applications plus the published context window above. Two tools can both be large language models underneath and still behave quite differently in practice, because the model is only one part of what a person actually experiences using the product.
Vendors also each publish their own position on whether a business’s own content trains their models. The stated position from major vendors is that it does not happen by default, in Microsoft’s case across Copilot’s own and third-party models it calls on, in Google’s case without explicit customer permission. That is a policy a business should read for itself in a vendor’s own documentation, not something to assume is the same everywhere.
None of this makes one large language model simply better than another in the abstract. A business already living inside Microsoft 365, with its email, calendar and documents in Outlook, Word and Teams, gets more out of Copilot’s grounding in that material than it gets from a stated context window it has no obvious use for. A business regularly working with a single very long document, a large contract or a technical manual, gets more practical benefit from Gemini’s published token figure than from a grounding feature built around a different ecosystem entirely. Choosing between them is a question about which product’s particular strengths line up with how a business actually works, not about which underlying model scores higher on an abstract benchmark.
Rollout is a Separate question from Capability
Buying a licence for a tool built on a large language model does not switch it on for everyone the moment the purchase is made. Admin-side controls decide who gets access to which feature and when, separately from what the model underneath is technically capable of. Microsoft’s Copilot Control System governs access per user and per app, with staged introduction of new features. Google’s admin console works through access groups, with some newer feature sets off by default until an administrator turns them on. A business rolling out a large language model based tool needs to plan that administrative step deliberately, rather than assuming a purchased licence and an available feature are the same thing on day one.
Where it gets things wrong
A large language model is trained to predict and generate language, not to check it. That single fact explains most of what goes wrong with it in practice. The model has no built-in sense of “I am not sure about this part.” It produces the next likely token whether or not that token is attached to anything real, and a sentence built that way can be completely wrong while reading exactly like a sentence that is completely right. This failure has a name, hallucination, and no amount of clever prompting removes it entirely, only reduces how often it happens.
The model’s knowledge also has a cut-off. It was trained on material available up to a particular point, and anything that happened after that point is simply not in it, unless the tool has been separately connected to a live source, a technique called retrieval augmented generation. Ask an ungrounded model about something recent and it will either say it does not know, the honest and useful answer, or answer confidently from older material as though nothing has changed, which is a hallucination with a specific, identifiable cause.
The model also reflects whatever the material it was trained on contained, including any skew in that material towards a particular viewpoint, group or language. That skew does not announce itself. A fluent, well-punctuated answer reads the same whether it is even-handed or not, which is exactly why a generated first draft needs a human check before it goes out under a business’s name, rather than being trusted on the strength of how confident it sounds.
Why the distinction from Generative AI matters in practice
The two terms get used almost interchangeably day to day, but they answer different questions, and getting them mixed up has a real cost. Generative AI is the umbrella term for anything that produces new content. A large language model is the specific kind of system under that umbrella built to do it with text. A business comparing ChatGPT, Copilot, Gemini and Claude is comparing large language models. A business comparing image tools for marketing material is comparing generative AI systems that are not language models at all, and the two comparisons do not use the same criteria.
That distinction also matters for anyone reading about AI regulation. The EU AI Act and most current rules are written to cover generative AI as a category, not language models specifically, because an image generator and a text assistant raise many of the same questions about accuracy, disclosure and oversight, even though only one of them is a language model. Reading “generative AI” in a regulation and mentally narrowing it to “chatbots” is a common and avoidable misreading, and it is one reason this site keeps the two terms separate rather than treating them as synonyms.
Where it is taught and where it shows up in an Audit
Large language models are the basis both for the courses this site runs and for the workflows an audit identifies, which is why the same vocabulary turns up in a training room and in an audit report. The core AI courses are built directly on the specific products a business already has, ChatGPT and Copilot most often, rather than taught as an abstract explanation of the technology underneath them.
An audit works the same way in reverse. Rather than assuming every process in a business wants the same tool, it looks at where a business’s actual workflows sit, which ones suit a large language model’s strength at producing a first draft or a summary, and which ones need the fixed, repeatable rule that automation provides instead. Knowing what a large language model actually is, and where its limits sit, is what makes that judgement possible in the first place, rather than reaching for the newest tool because it is the one currently getting attention.
FAQ
Questions about Large Language Model (LLM)
Is a Large Language Model the same thing as ChatGPT?
No. ChatGPT is a product built on top of a large language model. The model is the underlying technology that predicts and generates language. ChatGPT is one interface to a specific model, wrapped with a chat window, memory of the conversation and a set of features OpenAI chose to ship. Copilot, Gemini and Claude are the same relationship: a product name in front of a model doing the actual language work.
What is the difference between a Large Language Model and Generative AI?
Generative AI describes what a system does, which is produce new content rather than only classify or retrieve it. A large language model describes one particular kind of system built to do that with text specifically. Every large language model is a generative AI system. Not every generative AI system is a large language model, since an image generator is generative but was never trained on language at all.
What does a token and a Context window actually mean?
A token is the small chunk of text a model reads and produces one piece at a time, close to a word or a part of one. The context window is the total number of tokens a model can hold in view at once, the prompt, any attached document and its own answer combined. Google states a figure for this directly, one million tokens for Gemini, extended into tasks such as querying a large PDF in Drive without splitting it first. A bigger context window means a model can work across a longer document or a longer conversation without losing track of the earlier part of it.
Why does a Large Language Model sometimes give a wrong Answer with total confidence?
Because it is trained to predict a plausible next piece of text, not to check a fact against a source. A wrong statement and a correct one are produced by exactly the same process, so nothing in the tone gives away which is which. This failure has a name, hallucination, and it is the reason any output needs a human check before it goes out under a business's name.
Does a Large Language Model keep Learning from what we type into it?
Not in the way a person learns from a conversation. A model's core knowledge is fixed once training finishes, which is also why it carries a knowledge cut-off, a point after which it simply has not seen anything new. Major vendors state that business content typed into their tools is not used to train their models by default. That is a policy each vendor publishes and a business should read for itself rather than assume, not a technical property of how the model works.
Is a Large Language Model the same thing as an AI Agent?
No. A large language model answers one prompt at a time and stops. An AI agent uses a language model as one part of a system that can take several steps and act on other tools or systems with less oversight at each step. Every agent relies on a language model underneath it, but a language model on its own does not plan, act on a calendar or send an email by itself.
Why do ChatGPT, Copilot, Gemini and Claude all behave differently if they are all Large Language Models?
Because the underlying model is only part of the product. Each vendor trains its own model on its own material, tunes it differently, and wraps it in different features, Copilot's grounding in Microsoft Graph against a user's own emails and documents, Gemini's stated context window and native presence across Google Workspace, Claude's and ChatGPT's own separate chat products. Two tools can both be large language models and still produce noticeably different answers to the same prompt.
What it means for a Business
It is the basis both for the courses we teach and for the workflows an audit identifies, which is why the same vocabulary turns up in a training room and in an audit report.
Ready to Put This to Work?
Tell us where your team is with AI and we will tell you honestly what would make the biggest difference.