Data Protection
How personal data may be collected, used and stored wherever an AI system touches it. Taught as a set of practical choices rather than as law, and checked alongside the EU AI Act at an audit's Identify step.

What Data Protection means here
Data protection is how personal data may be collected, used and stored wherever an AI system touches it. On this site it is taught as a set of practical choices and not as law. The regulation behind it is GDPR, and this page is the practical companion to that one.
Where the GDPR page explains the rules, this page covers the decisions a business makes because of them. Where does a model run? What is removed before text goes in? What may staff enter, and into which tool? It is not legal advice, and the site says no audit substitutes for advice from a qualified lawyer.
Choices, not a Checklist of Law
The site’s courses do not teach data protection as a lecture on a statute. The HR course closes on it as a run of concrete choices. It covers cloud versus local hosting, running a model with the wifi switched off, anonymisation done live and offline, and where the regulation puts responsibility for a screening decision.
The reason is practical. A person who has watched a model run with the network off understands what local means in a way a definition does not give. The stated outcome of that course is that participants anonymise data locally and hold a defensible position on CV screening. That is a skill, not a fact to memorise.
Where the Data goes
The first practical question about any tool is where the data goes. A hosted model is called over an API or through a product built on one, and the data is processed on the provider’s infrastructure under the terms of whichever agreement the account is on. A self-hosted model runs on infrastructure the organisation controls, and nothing sent to it leaves that environment.
The site’s reviews put the trade-off plainly. Self-hosting buys direct control over where data physically sits and who can reach it, and costs hardware and the people to run it. A hosted model gives capability and speed without infrastructure, and puts the data path on the provider’s side. The local model against Claude review and the cloud against on-premise review set out both.
Neither location is automatically Compliant
Two shortcuts are wrong in opposite directions. One says a hosted service is the risky option. The other says running a model on your own hardware makes a business compliant. The site’s reviews say obligations under the EU AI Act and GDPR attach to what a system does with data, not to where it is hosted. A poorly governed local deployment can breach either as easily as a poorly governed hosted one.
What changes is who is directly accountable for the data path and how much of it the organisation can see and control. On-premise makes it more straightforward to show what happens to data at each step. A cloud deployment’s position depends on the provider’s documented terms being read, checked and kept on file. A hybrid, splitting workloads by sensitivity, is a standing third option, decided workload by workload against documented requirements.
What the vendors state
The site’s tool pages report each vendor’s own statements and do not verify them. Anthropic states that inputs and outputs from standard accounts may be used to train future models unless the user opts out, and that Business and Enterprise accounts are governed by their own customer agreement. OpenAI states that its business tiers and its API are excluded from model training by default unless an organisation opts in, and offers data residency in the EEA and Switzerland on some plans.
Google states that Workspace, including whatever Gemini processes inside it, is not used to train its generative AI models without the customer’s explicit permission, and offers an EU data region for those integrations. Microsoft states that prompts, responses and data accessed through Microsoft Graph are not used to train the foundation models behind Copilot, and that traffic for EU users stays within the EU Data Boundary. Each of those pages sets out the qualifications the vendor itself notes.
Read the account you will use
The point that repeats across those pages is the account. Two people in the same company can use the same assistant under different terms, depending on which sign-in they use. A standard consumer account and a business account are different agreements. The Microsoft page notes that its protection depends on a work sign-in, and the Gemini page notes the consumer app is a different setting.
So the site’s advice is to read the terms for the product and plan the business will actually use, and to decide what staff may enter. It also says the documented policy for a developer platform and a business tier’s customer agreement are separate documents, and neither should be assumed to carry over to the other.
What Certificates do and do not say
The Gemini review lists Google’s certifications, and the ChatGPT page lists OpenAI’s. The Gemini page observes that OpenAI’s cover largely the same ground, so security written down on paper is rarely what decides the choice. That is worth holding on to. A certificate is a statement about how a vendor runs its own systems.
It does not tell a business whether its own use is fine. The EU AI Act page makes the same point about a vendor’s compliance statement: it says what the vendor commits to, and does not say whether the business’s use is compliant. The business still has to decide what goes in, and be able to show how it decided.
What staff may enter
Most exposure begins with a person and a tool. The site’s answer is a one-page rule, drawn by kind of data and not by tool. The first part lists what may never be entered into any AI tool: client names and contact details, contracts, HR records, credentials and anything under a confidentiality agreement. The second lists what is fine in the approved tool, such as drafting from your own notes, summarising public documents and editing text with the names removed. The third says anything else goes to one named person before it is entered.
The banning article adds two firm cases. A specific tool can be unacceptable for a specific kind of data, for example a consumer service whose terms let it retain and reuse what you enter, where client material sits under confidentiality. And where GDPR applies to the personal data involved and no approved tool with suitable terms exists yet, the honest answer is not to use AI on that data until one does. It also notes that personal data entered into an unassessed tool can breach a business’s GDPR duties.
Anonymisation
Removing what identifies a person before text goes into a tool is one of the practical choices the site teaches. The HR course does it live and offline. The admin and document workflows course works secure document translation through three steps: classification, automated anonymisation and private translation.
The course is candid about the limit. It lists that use case under the heading of why it has no clean answer today. That is a useful model for the whole subject. Anonymisation is a step in a process and not a guarantee, and it belongs beside the decision about where the model runs and who can see the output. The personalisation course closes on the safest version of the rule: public, anonymised, generic or invented content only.
Agents and Knowledge bases
An assistant that reads documents or reaches systems is a decision about data. The agents course includes governance and security essentials, and its build covers user authentication and controlled topics for high-risk policies. What goes into an agent’s knowledge base is a decision in its own right, so the line is drawn by kind of data and not by tool.
The vendors state that business content is not used to train their models by default, and that is a policy each vendor publishes and a business should read for itself. The AI agent page covers the wider question of what an agent is allowed to reach.
A worked example
This is an illustration, not a real client. A small HR team wants to use AI to tidy candidate CVs before a hiring manager reads them. The CVs contain names, contact details and work histories, so they are personal data. The team asks three practical questions. Does the text need to go to a hosted tool at all, or can a local model do the job? If it does go out, can the identifying details be removed first? And which account will the team use, and what do that account’s terms say?
The team settles on removing names and contact details locally before anything is entered, uses the business account and not a personal one, and keeps a note of how it decided. The screening decision itself stays with a person, because the HR course says human oversight is the part that can be held responsible. Nothing here needed a specialist. It needed the questions asked before the tool was used.
In Training and in an Audit
Data protection is built into every foundational programme. The all-staff briefing covers data privacy and information sensitivity, approved against third-party tools, what can go wrong and how to report an incident. Each ChatGPT course opens with security and the use of private data, along with safety rules for working in third-party tools, before any tool is opened.
In an audit it is checked at the Identify step, beside the EU AI Act. The audit article says where a tool would touch personal data, it is checked under data protection rules before the tool is recommended, because a problem found after roll-out costs far more to unwind. It also asks where a specific piece of information physically sits and who is able to open it. The AI audit page states the limits.
Mix-ups worth avoiding
Four confusions come up often enough to name. The first is treating data protection and GDPR as the same thing, when one is the set of practical choices and the other is the regulation behind them. The second is assuming a hosting choice settles it, in either direction.
The third is treating a vendor’s statement or certificate as the end of the matter, when it describes how the vendor runs its systems and not whether your use is fine. The fourth is treating anonymisation as a guarantee. It is a step, and it has to sit inside a process.
Where it goes wrong
One failure is checking too late, when a tool is already in daily use and someone asks afterwards what happened to the data. The site’s answer is to check before the tool is recommended. Another is assuming one policy covers every account, when a standard account and a business account sit under different terms.
A third is a rule staff cannot apply. The banning article says a rule people understand gets followed in cases it did not anticipate, so staff need enough AI literacy to judge a case the page never covered. A fourth is a rule with nobody responsible for it, which is a matter for AI governance.
Where it is taught
Data protection runs through the foundational programmes. The HR course closes on it, the admin and document workflows course works it through secure translation, and the all-staff briefing covers it for a whole organisation at once. The rest sit in the wider training catalogue.
The regulation itself is on the GDPR page, and the wider rules an organisation writes on top of it are on the AI governance page.
FAQ
Questions about Data Protection
What is Data Protection in AI?
How personal data may be collected, used and stored wherever an AI system touches it. The site teaches it as a set of practical choices and not as law, and checks it alongside the EU AI Act at an audit's Identify step. GDPR is the regulation that sets the rules.
What is the difference between Data Protection and GDPR?
GDPR is the EU regulation that covers how personal data may be collected, used and stored. Data protection, as the site uses the term, is the set of practical choices a business makes on the back of it: where a model runs, what is anonymised, and what staff may enter into a tool.
What should staff never enter into an AI Tool?
The site's one-page rule lists client names and contact details, contracts, HR records, credentials and anything under a confidentiality agreement. The line is drawn by kind of data and not by tool, and anything not on the list goes to one named person before it is entered.
Is our Data used to train the AI?
It depends on the vendor and on the account. The vendors state that their business products are not used for model training by default, and that standard consumer accounts can be under different terms. Those are the vendors' own statements. Read the agreement for the account your staff will actually use, because two people in one company can sit under different terms.
Does keeping Data in the EU solve it?
Not by itself. Some vendors offer EU or EEA data residency on certain plans, and the site reports those statements, but residency is one setting and not the whole of data protection. The Act and GDPR attach to what a system does with data, wherever the software runs.
Is a Local Model Safer than a Cloud one?
It gives a business more direct control over where data sits and who can reach it, and it also puts the maintenance and accountability on the business. Neither option is automatically the more compliant one. A hybrid split by sensitivity is a standing third option, decided workload by workload.
Local LLM vs Claude: Which fits your Business?Cloud AI vs On-Premise AI: Which fits your Business?
Is anonymising Data enough?
It is one of the practical choices the site teaches, and the HR course does it live and offline. Secure document translation, in the admin course, is worked through as classification, automated anonymisation and private translation, and the course says that use case has no clean answer today. Anonymisation is a step in a process and not a guarantee.
AI Essentials for Human resourcesAI Essentials for admin & Document Workflows
Does an AI Audit check Data Protection?
Yes. Where a tool would touch personal data, the audit checks it under data protection rules before the tool is recommended, not after it is in use. It is a diagnostic and not a legal opinion, and no audit substitutes for advice from a qualified lawyer.
What it means for a Business
It closes one course as a set of practical choices: cloud versus local hosting, running a model with the wifi switched off, anonymisation done live, and where the regulation puts responsibility for a screening decision.
Ready to Put This to Work?
Tell us where your team is with AI and we will tell you honestly what would make the biggest difference.