Automating a Task vs Automating a Decision
Only 5% of jobs are fully automatable against 45% of tasks. That gap is the whole argument for sorting work onto a ladder instead of a binary.
11 min readBy Alex Frolo

On this page
A task automated well saves time. A decision automated without being scoped for it produces a result nobody signed off on, and the two get confused constantly, because they usually sit inside the same workflow.
This article sets out why that confusion happens, walks through a four-level test that sorts a piece of work onto the right side of the line, and looks at what the accuracy numbers actually say about where AI supports a decision well and where it does not.
Key takeaways
- Industry-wide, AI can automate up to 45% of work activities, but only 5% of jobs are fully automatable. The gap between those two figures is the task and decision distinction, not a rounding error.
- AI decision accuracy drops to around 65% in ambiguous, judgement-heavy situations, against roughly 85% for experienced humans in the same conditions.
- On structured, data-rich predictions such as candidate shortlisting, AI reaches about 75 to 80% accuracy against 85 to 90% for experienced humans, close enough to support the call, not replace it.
- A four-level test sorts any piece of work: Assist, Recommend, Execute, or Never delegate, based on volume, consequence, reversibility, how clear the rules are, and whether the decision needs legitimacy a model cannot supply.
- Hiring, dismissal and anything with legal, safety or rights implications sit permanently at Never delegate, whatever the technology in front of them can do.
- A well-scoped AI workflow is a workflow where every step inside it has already been sorted onto that ladder, not approved as one undifferentiated block.
Why the two get treated as one thing
“Automate the process” usually means automating a chain of tasks with one or two decisions buried inside it, and the whole chain gets approved together as if every step carried the same risk. Formatting a report, pulling figures into a template, routing a ticket to the right queue: those are tasks, checkable against a right answer, and safe to hand over completely once the rules are set. Deciding which candidate gets a callback, which claim gets flagged as fraud, which client gets an exception to a policy: those are decisions, and getting one wrong costs something a checked task never costs, because a person somewhere has to explain the outcome.
The scale of that gap is worth putting a number on. Industry-wide, AI can automate up to 45% of work activities, but only 5% of jobs are fully automatable, according to analysis compiled by SkillSeek. The 45% figure describes how much of the average job is task-shaped, made of steps that can be checked against a fixed standard. The 5% figure describes how rarely an entire role, decisions included, can be handed over without a person still answering for the result. Most of the distance between those two numbers is the task and decision distinction this article is about, not a measurement error or a gap that will simply close as the models improve.
That distinction matters because a manager approving “AI for the intake process” is rarely approving a single thing. A real intake process is a chain: collect the form, check it against a checklist, flag anything unusual, and decide what happens next. The first two steps are tasks. The third step might be a task if the rules are genuinely fixed, or it might be the decision the whole process exists to make. Treating the chain as one approval is how a business ends up with an AI system quietly making calls nobody scoped it to make, discovered only once a result looks wrong and somebody asks who signed off on it.
Why the accuracy gap is the reason, not an afterthought
The instinct to treat automation as a single dial, more automated against less automated, ignores that AI’s own accuracy is not constant across the two categories. AI decision accuracy drops to around 65% in ambiguous, judgement-heavy situations, against roughly 85% for experienced humans working the same call, per a meta-analysis cited by SkillSeek. That is not a small gap, and it sits precisely in ambiguous, judgement-heavy territory, which is exactly where a decision, rather than a task, tends to live. A task rarely asks a system to weigh conflicting factors with no fixed answer; a decision almost always does.
On structured, data-rich predictions, the picture looks better but still favours the human. AI reaches about 75 to 80% accuracy on calls like candidate shortlisting, against 85 to 90% for experienced humans doing the same job. That gap is close enough for the AI’s output to be genuinely useful as a second opinion, ranking or flagging cases for a person to look at first. It is not close enough to hand the call over and stop checking it, and treating a 75 to 80% system as though it clears the same bar as an 85 to 90% one is where the confidence outruns the evidence.
Neither number is a verdict on the technology overall. It is a description of where AI’s own error rate sits relative to a human’s, at each end of the ambiguity scale, and that description is exactly the information a business needs before deciding how much authority to hand over. A system that is 20 points behind a human on the hardest calls and 10 points behind on the easiest ones is not one dial, it is two different tools doing two different jobs, and the accuracy numbers are the evidence for treating them that way rather than a reason to distrust automation across the board.
A four-level test, not a binary
Harvard Business School Online frames the underlying design question well: the real choice is not whether something can be automated, it is whether AI should replace judgement or support it, sorted by how often the work happens and how much risk and value ride on getting it right. A practical way to apply that is a four-level ladder, adapted from a decision-delegation framework AI strategist Kieran Gilmurray has laid out for boards weighing exactly this question.
Assist
AI drafts, gathers, compares options. The person decides. Fits ambiguous, strategic calls.
Recommend
AI ranks or suggests, a person stays accountable. Fits structured, data-rich calls like fraud alerts or shortlisting.
Execute
AI acts inside clear rules, unsupervised. Fits high-volume, low-consequence, reversible steps.
Never delegate
Hiring, dismissal, anything with legal, safety or rights implications. Stays human regardless of capability.
The rungs are sorted by five questions, not by what the technology can technically attempt: how often the work happens, what a wrong outcome costs, whether it can be undone, how clear the governing rules already are, and whether the outcome needs a kind of legitimacy a model cannot hold. A password reset scores low on every risk factor and high on volume, which is why it sits at Execute. A dismissal scores the opposite on every factor that matters, which is why no amount of model capability moves it off Never delegate.
The same five questions explain why two pieces of work that look similar on the surface end up on different rungs. A password reset and a candidate shortlist are both, in a narrow sense, “deciding something about a request.” The password reset is reversible in seconds, happens thousands of times a day, and is governed by a rule with no ambiguity in it: the request either matches the account or it does not. Candidate shortlisting happens far less often, is much harder to undo once a candidate has been told no, and is judged against criteria that are never perfectly fixed in advance. Run both through the same test and they land on different rungs for reasons that have nothing to do with which one AI is technically capable of doing.
That picture is the whole argument compressed into two bars. The password reset climbs all the way to Execute, because every question the ladder asks about it comes back low-risk. Candidate shortlisting stops at Recommend, one rung lower, because it scores high enough on consequence and ambiguity that a person has to stay accountable for the final call, however good the AI’s ranking is underneath it.
What this means for scoping an actual workflow
The practical use of the ladder is not philosophical, it is a scoping exercise. Every real AI workflow under consideration is a chain of individual steps, and each step belongs on its own rung: the intake might be Execute, the triage might be Recommend, and the final call might be Assist at best, or Never delegate outright. Treating the whole chain as one automation decision is how a task-shaped process quietly starts making decision-shaped calls nobody scoped it to make, and it is usually discovered only after something has already gone out the door.
A well-scoped AI workflow covers what that scoping looks like once it is done properly, and why the return concentrates where it is. The pattern that shows up there is the same one this article’s ladder predicts: the steps that were correctly sorted onto Execute deliver a fast, measurable return, because nobody is second-guessing a step that never should have needed a person watching it. The steps left at Assist or Recommend take longer to show a return, because the return there is a better-informed person making the call faster, not a person removed from the loop altogether. Confusing the two, expecting an Assist-level step to deliver Execute-level speed, is a common reason a workflow gets judged a disappointment when the actual fault was in how it was scoped.
This is also where an audit earns its keep, rather than being a compliance exercise bolted onto the end of a project. Sorting a process onto the ladder before anything is built means candidate shortlisting type steps get scoped as Recommend from day one, with a named reviewer built into the process, instead of drifting towards Execute later because nobody revisited the original design once the tool proved itself reliable on the easy cases. What an AI audit actually checks covers this in more detail: an audit is largely this same ladder, applied to processes that are already running rather than ones still on a whiteboard, and it is often the first time anyone has looked at the whole chain rather than the one step that prompted the review.
The timing matters as much as the sorting itself. What predicts whether an AI project reaches production found that projects stalling after a promising pilot were disproportionately ones where a Recommend-level step had quietly been treated as Execute-level once the demo looked good, and someone senior caught it before it shipped rather than after. Catching that at the scoping stage costs an afternoon. Catching it after launch costs a rebuild, and sometimes an apology to whoever the wrong call affected.
None of this holds unless the people running the workflow day to day actually know which rung each step sits on, rather than that judgement living in one manager’s head. How to brief a team on AI covers what a team needs to arrive with before that conversation happens, so the ladder is not news to anyone in the room when a step is being sorted.
Where governance and workflow design meet
Governance and workflow design are frequently treated as separate disciplines, one for the compliance file and one for the delivery team, and that separation is exactly where processes drift off the rung they were scoped for. A written AI governance policy is, among other things, a record of which rung each automated process sits on and who signed off on putting it there. Without that record, a business finds out its process crossed from Execute into Never-delegate territory only after something goes wrong, rather than at the point the workflow was designed. Good workflow design already asks the same five questions a governance review asks; keeping them in separate documents, reviewed by separate people on separate schedules, is how a process ends up governed on paper and misclassified in production.
The legal backdrop makes the stakes concrete rather than abstract. The EU AI Act draws its own line around what counts as a high-risk AI system, and the categories it names, employment decisions among them, map closely onto the Never delegate rung this article has been describing. That overlap is not a coincidence. Both the legal framework and the practical test are answering the same underlying question: which decisions carry enough consequence and enough need for accountability that no amount of model accuracy changes the answer. A business that has already sorted its workflows onto this ladder is, in practice, most of the way to a defensible position on the regulatory question too, because the same five factors that decide where a step sits on the ladder are what a regulator will ask about first.
None of this is an argument against automation. It is an argument for measuring the right thing when deciding how far to take it. A process judged only on how many business outcomes it improves, without anyone asking which of those outcomes were tasks completed faster and which were decisions made with less human judgement than the situation called for, is being measured with half the information a proper review needs.
What follows
Before automating a process, list its individual steps and sort each one onto the ladder rather than approving the chain as a whole. Anything high-volume, reversible and governed by clear rules can move to Execute. Anything ambiguous or judgement-heavy stays at Assist or Recommend, with a named person accountable for the call. Anything touching hiring, dismissal or a legal or safety obligation stays off the ladder entirely, however capable the tool in front of it looks.
If that sorting exercise has not been done for the processes already running in your own business, a scoping conversation is where it starts, working from the same five questions this article has used throughout, applied to what is actually running today rather than to a hypothetical.
FAQ
Questions we get asked
What is the difference between automating a task and automating a decision?
A task is a repeatable step with a right answer that can be checked: formatting a document, routing a ticket, resetting a password. A decision carries judgement, and getting it wrong has a cost a checked step does not: who gets hired, whose loan is declined, whose access is revoked. Automating the task removes drudgery. Automating the decision transfers accountability, and that transfer needs to be deliberate rather than a side effect of automating the task sitting next to it in the same process.
What does human in the loop actually mean in practice?
It means a named person reviews the AI's output before it takes effect, for any decision where being wrong is costly or hard to reverse. It is not a checkbox on a settings page. A human in the loop who rubber-stamps every recommendation without reading it provides no real oversight, which is why the review step needs a genuine point of friction, not just a signature.
What is the difference between automation and augmentation?
Automation lets a system act without a person in the loop, for high-volume, low-consequence, reversible steps such as filtering spam or routing a routine request. Augmentation keeps a person deciding, with AI supplying the drafting, comparison or ranking that speeds the judgement up rather than replacing it. Harvard Business School Online frames the design question the same way: not whether AI can act, but whether it should replace judgement or support it.
Which decisions should never be delegated to AI?
Hiring, dismissal, and anything carrying legal, safety or rights implications, whatever the system could technically do. Those decisions need legitimacy and accountability that a model cannot hold, so the test is not capability, it is whether a person has to answer for the outcome. If the answer is yes, the decision stays human even where an AI tool produces a defensible recommendation.
How accurate is AI compared to humans at making decisions?
It depends heavily on how ambiguous the situation is. AI decision accuracy drops to around 65% in ambiguous, judgement-heavy situations, against roughly 85% for experienced humans in the same conditions. On structured, data-rich predictions such as candidate shortlisting, the gap narrows to about 75 to 80% for AI against 85 to 90% for experienced humans, close enough to support a decision without being trusted to make it alone.
What is a reversible versus an irreversible decision in AI automation?
A reversible decision can be undone at low cost if it turns out wrong: a routine expense approval, a routing choice, a password reset. An irreversible one cannot: a dismissal, a declined application, a public commitment made in a client's name. Reversibility is one of the clearest tests for how much authority a system should be given, alongside volume, consequence and how clear the rules already are.
How do I decide how much authority to give AI in a specific workflow?
Sort the workflow by five questions. How often does it happen, what does a wrong outcome cost, can it be undone, how clear are the rules already, and does the decision need legitimacy a model cannot supply. High-volume, low-consequence, reversible steps with clear rules can run on their own. Anything scoring the opposite on those five stays with a person, whatever the AI in front of it can technically do.
