Jul 26, 202610 min readDarek Ambroziak

    Where Does AI Fit Without Replacing Human Judgment?

    AI fits in the coordination layer: gathering evidence, drafting options, monitoring processes. Human judgment stays where a decision is irreversible, contested, or attributable. Three tests, one decision table, and a 90-day plan for drawing the line.

    Editorial illustration of a dividing line between warm organic forms representing human judgment and a structured data grid representing AI as a coordination layer.

    AI fits in the coordination layer: gathering evidence, drafting options, monitoring processes, flagging exceptions. Human judgment stays where a decision is irreversible, contested, or attributable to someone. The dividing line is not task difficulty or seniority. It is whether the output is material a person evaluates — or a verdict a person is asked to rubber-stamp.

    TL;DR

    • AI belongs in the coordination layer. It moves, drafts, sorts and watches. Judgment is the act of committing the organization, and it stays with people because accountability stays with people.
    • The evidence supports the split. Across 106 experiments, combinations of humans and AI lost performance on decision tasks and gained on content-creation tasks (Vaccaro, Almaatouq & Malone, 2024).
    • Three tests draw the line: is the outcome reversible, is it contested, is it attributable? One 'yes' keeps judgment with a person.
    • 'Human in the loop' is not a safeguard by default. EU AI Act Article 14 requires overseers who can understand, disregard and stop the system. A reviewer without time, information or authority is decoration.
    • Judgment does not disappear as autonomy rises — it relocates, from the transaction to the design of the system that handles transactions. Organizations that miss the relocation lose judgment by accident.

    What is human judgment, and what is it not?

    Human judgment is the act of committing an organization to a course of action that is irreversible, contested, or attributable — under incomplete information. Retrieval, classification, ranking, calculation and drafting are not judgment. They are inputs to it.

    That definition matters because most organizations draw the line in the wrong place. They define judgment by seniority ("the director signs it") or by difficulty ("this is complex, so a person should do it"). Both are wrong.

    Reconciling an invoice against a 4,000-line contract is hard and is not judgment — it is pattern work at volume. Telling a long-standing customer that their claim is denied looks trivial and is judgment — it commits the firm, it is contestable, and someone's name is on it.

    What does the research say about putting AI inside decisions?

    The strongest available evidence is a preregistered systematic review of 106 experimental studies reporting 370 effect sizes, published in 'Nature Human Behaviour' (Vaccaro, Almaatouq & Malone, 2024). Every study compared humans alone, AI alone, and the two combined.

    Three findings should shape where you place AI.

    First, combining is not automatically better. On average, combinations of humans and AI performed significantly worse than the best of humans or AI alone (Hedges' g = −0.23; 95% CI −0.39 to −0.07).

    Second, the losses cluster in decisions. The study found performance losses on tasks that involved making decisions and significantly greater gains on tasks that involved creating content. That is the answer to the title question, stated empirically: AI fits where the work is generative, and fits badly where the work is a verdict.

    Third — and this is the design rule — the direction of the effect depends on who was better to begin with. When humans outperformed AI alone, the combination produced gains. When the AI outperformed humans alone, the combination produced losses.

    Read that plainly: putting a person in front of a system that is better than them degrades the result. Putting a system in front of a person who is better than it improves the result. You do not staple the two together and assume the sum beats the maximum.

    Why doesn't 'human in the loop' protect judgment on its own?

    Because oversight without capacity is theatre, and the regulator has already said so.

    EU AI Act Article 14 requires that high-risk systems be designed so that they can be effectively overseen by natural persons. Two sub-provisions matter here. Article 14(4)(b) requires that overseers remain aware of the tendency to over-rely on system output — the law names automation bias explicitly. Article 14(4)(d) requires that they be able to decide not to use the system, or to disregard, override or reverse its output. See our deep dive on human in the loop AI for the operating model.

    If your reviewer cannot practically say no, you do not have oversight. You have a signature.

    The failure runs in both directions. People over-rely on systems they should question, and they abandon systems they should trust: after watching an algorithm make an error, people lose confidence in it faster than they lose confidence in a human who makes the same error (Dietvorst, Simmons & Massey, 2015). The design goal is calibrated reliance — trust matched to demonstrated reliability — not maximum trust (Lee & See, 2004). See trust calibration in human-AI collaboration.

    And accountability does not transfer. In 'Moffatt v. Air Canada' (2024 BCCRT 149), the airline argued that its chatbot was a separate legal entity responsible for its own actions. The tribunal rejected that outright: the company was responsible for what its system told a customer. Deploying a system does not move the liability onto the system.

    Which tasks should AI take, and which stay with people?

    Reading the evidence — retrieval and summarization at volume — sits fully with AI, which can read what nobody had capacity to read. What stays with people is deciding what counts as evidence. This is Augment on the AI Value Ladder.

    Producing options is content creation, the strongest measured territory for AI, and belongs to AI fully. What stays with people is deciding which options are admissible. Assist to Augment.

    Routine classification that is reversible and uncontested is pattern matching. Let AI run it with sampled review, and keep thresholds and exception rules with people. Automate.

    Scoring a person — hiring, credit, performance — is ranking under contested values. AI can only assemble evidence; the decision and its written justification stay with a named person. Assist only.

    Committing the organization — pricing, safety, care, legal position — is irreversible and attributable. AI helps with preparation and second opinion; the decision stays with a person, without exception. Assist.

    Watching the live system is anomaly detection and belongs fully to AI. What counts as an anomaly worth acting on stays with people. Automate to Agentic.

    How do the four rungs of the AI Value Ladder change who decides?

    The AI Value Ladder describes four levels of autonomy. Judgment behaves differently at each.

    • Assist — AI drafts, a person sends. Judgment is untouched; only the cost of preparing it falls.
    • Augment — AI does what nobody had the capacity to do, such as reading every complaint filed this quarter. Judgment expands, because it now rests on more evidence.
    • Automate — AI runs the process and a person handles exceptions. Judgment moves upstream, into threshold and escalation design.
    • Agentic — AI plans, uses tools and acts inside guardrails. Judgment moves further upstream still, into guardrail design and stop conditions.

    Judgment does not disappear as you climb the ladder. It relocates. At rung one, judgment sits in the transaction. By rung four, it sits in the design of the system that processes ten thousand transactions. Organizations that fail to notice the relocation do not eliminate judgment — they lose it by accident, because nobody was assigned the upstream decision. See Human-in-the-Loop vs Human-on-the-Loop for the levels-of-autonomy split.

    Why can AI never be a party to a decision — only a layer beneath it?

    Collaboration is a social process. It happens between people, and it requires three conditions: aligned goals, compatible attitudes, and mutual knowledge of competencies (Wekselberg framework).

    AI meets none of them. It has no goals of its own to align. It has no attitudes. It has no knowledge of what your colleagues are actually capable of. AI is a coordination layer — it moves information between people faster and at greater volume than they could move it themselves. People collaborate on top of that layer. The layer is not a participant.

    This is an accountability point, not a semantic one. A decision requires a party who can be held to it, and only a person can be held to anything.

    It is also why the line between AI and judgment is drawn by management, not by technology. Where goals are not aligned across levels, AI simply accelerates work toward conflicting goals. That is visible in the numbers: roughly 95% of enterprise generative AI pilots produce no measurable P&L impact (MIT Project NANDA, 2025), and around 70% of the value in an AI initiative comes from people and process change rather than algorithms and technology (BCG, 2024). Among 100+ Polish managers, the two weakest of six AI maturity dimensions were change management (2.42/5) and strategy and governance (2.51/5) — the exact dimensions that decide who decides (AI_Managers™ 2026).

    How do you draw the line in your own organization in 90 days?

    • Days 1–15 — Inventory decisions, not tasks. List every recurring decision in one process, with the person accountable for each. Most organizations have never written this down, which is why AI gets deployed against tasks and quietly absorbs decisions.
    • Days 10–25 — Apply the three tests. For each decision: is the outcome reversible, is it contested, is it attributable? Three 'no' answers mean it is not judgment — automate it. One 'yes' keeps a named person on it.
    • Days 20–40 — Assign a rung, not a tool. Place each decision on the AI Value Ladder based on the cost of being wrong, then choose technology. Do the reverse and the tool will choose the autonomy level for you.
    • Days 30–55 — Equip the overseer with three things. Time to review, information to review with, and the authority to override without career cost. Article 14(4)(d) assumes all three; most deployments provide none.
    • Days 45–75 — Instrument the loop. Track override rate, escalation rate, and time-to-decision. An override rate near zero is a warning, not a success: it usually means the reviewer has stopped reviewing.
    • Days 60–90 — Move the result into the P&L. Report the decision quality change, not the licence count or the adoption rate.

    These map to three RECODE Method dimensions: Redesign Work (steps 1–3), Establish Ownership (step 4), and Operationalize Value (steps 5–6). See the full RECODE Method framework.

    The regulatory clock supports doing this deliberately. Following final Council approval of the Digital Omnibus on 29 June 2026, high-risk obligations under Annex III of the EU AI Act apply from 2 December 2027, and Annex I obligations from 2 August 2028 (Council of the EU, 2026). That is time to design oversight properly, not time to defer it.

    FAQ

    What is human judgment in an AI context?

    Human judgment is the act of committing an organization to a course of action that is irreversible, contested, or attributable to a named person, under incomplete information. Classification, retrieval, ranking and drafting are not judgment — they produce the material judgment operates on. AI belongs on the material side of that line.

    Where does AI add the most value without touching judgment?

    In content creation and evidence work: drafting, summarizing, translating, and reading volumes of material nobody had capacity to read. The 2024 'Nature Human Behaviour' meta-analysis found significant gains for combinations of humans and AI on content-creation tasks and losses on decision tasks. Put AI where it generates, not where it concludes.

    Does the EU AI Act require a human to make the final decision?

    Not universally. Article 14 requires that high-risk systems be effectively overseen by natural persons who can understand the system's limits, remain aware of automation bias, and disregard, override or reverse its output. It regulates the capacity to intervene rather than mandating that a person press every button.

    Is a low override rate a sign the AI is working?

    Usually the opposite. An override rate near zero generally means the reviewer has stopped reviewing — the rubber-stamp failure mode Article 14(4)(b) is written against. Healthy oversight produces a stable, non-trivial rate of disagreement, and each override should be logged as training signal for both the model and the process.

    Doesn't this mean people and AI end up collaborating?

    No. There is no such thing as human-AI collaboration. Collaboration is a social process between people and requires aligned goals, compatible attitudes and mutual knowledge of competencies — conditions AI cannot meet. People collaborate with each other; AI is a coordination layer that supports, augments and accelerates their work.

    Sources

    • Vaccaro, M., Almaatouq, A., & Malone, T. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8, 2293–2303.
    • Regulation (EU) 2024/1689 (EU AI Act), Article 14 — Human oversight.
    • Council of the EU (2026). Final approval of the Digital Omnibus on AI, 29 June 2026.
    • Moffatt v. Air Canada, 2024 BCCRT 149 (British Columbia Civil Resolution Tribunal, 14 February 2024).
    • Dietvorst, B. J., Simmons, J. P., & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1), 114–126.
    • Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50–80.
    • MIT Project NANDA (2025). The GenAI Divide: State of AI in Business 2025.
    • Boston Consulting Group (2024). The 10-20-70 rule for AI value creation.
    • AI_Managers™ 2026 (January 2026). AI maturity study, 100+ Polish managers.
    • Wekselberg, V., & Wasilewski, J. (2023). Cooperation, collaboration, coordination, groupthink. Difin.
    #AI Governance#Human Oversight#EU AI Act#Decision Making#AI Value Ladder
    Share