H2H vs H2M: Why Cross-Functional Collaboration Is the Real Innovation Engine
Innovation comes from people across functions working on aligned goals — not from another tool. Why H2M stalls without H2H, and the four mechanisms that turn diverse functions into one engine.
Most companies are pouring money into AI tools and wondering why the productivity curve stays flat. BCG's 2025 research is blunt about it: only 5% of organisations are 'future-built,' and nearly 90% of AI leaders expect most of their AI value to come from reshaping cross-functional business processes — not from the tools themselves.
That gap has a name. Human-to-machine (H2M) collaboration — prompting, verifying, delegating to AI — is the visible layer. Human-to-human (H2H) collaboration — how distinct functional goals are aligned, how boundaries are spanned, how disagreement is made productive — is the scaffolding underneath. Skip H2H and AI just makes your silos faster. (For the short working definition, see What is H2H vs H2M readiness?.)
Coordination is not collaboration
Most organisations confuse the two. Coordination is keeping parallel work from colliding. Collaboration is making distinct goals compatible across individual, team, department, and organisation — the textbook signature of teams that out-innovate. Without that alignment, every AI output gets re-litigated by hand, every handoff piles up at the same bottleneck, and the license fees keep auto-renewing.
The four mechanisms that actually move the needle
Our work — and the underlying research — converges on four mechanisms that turn diverse functions into one innovation engine. AI then amplifies the work; it does not replace it.
- Aligned goals across levels. Not 'common' goals — compatible distinct goals across individual, team, department, and organisation. This is where most AI rollouts quietly fail.
- Boundary spanning by design. Teams that scan markets, coordinate across functions, and engage leadership secure more resources and produce more innovation than inward-looking ones (Ancona & Caldwell; ~5.4× advantage).
- Cognitive diversity, productively used. Different functional lenses generate friction. Channelled well, that friction is the raw material of breakthroughs.
- Psychological safety. AI value leaks the moment people are afraid to flag a wrong output, a brittle workflow, or a bad assumption.
Why weak H2H breaks H2M
Dell'Acqua and colleagues (HBS / BCG) found consultants using GPT-4 on in-frontier tasks produced output ~40% higher in quality and worked 25% faster — but only when human judgment guided where AI should help. Hand the same tool to a team without aligned goals or safety to challenge outputs, and the lift disappears. The model isn't the variable. The H2H around it is.
Three diagnostic questions
- Can two team members from different functions independently rate the same AI output and agree within 10%? If not, you have an H2H quality-standard problem, not an H2M problem.
- When AI output is rejected, does the feedback flow back into the workflow — or die in someone's head? If it dies, you have an H2H learning-loop problem.
- Who decides when to escalate from AI to human? If the answer is 'it depends,' you have an H2H authority problem — and AI will only make it more expensive.
Calibrated trust in practice: three worked examples
The examples below are composite scenarios based on patterns described in the research above. They show how the same AI tool gives very different results depending on the human work around it.
Example 1: Marketing and legal review the same AI draft
Marketing uses AI to draft campaign copy; legal rejects half of it. Without H2H alignment, each rejection turns into an email thread, and marketing starts to accept or reject AI drafts on gut feeling. With a shared, written acceptance standard (what claims need a source, which words are off-limits), both functions rate drafts the same way. People trust the tool where it has proven reliable (tone, structure) and check it where it tends to fail (product claims). That is calibrated trust: it depends on the task, not on whether someone likes the tool.
Example 2: A product team spans a boundary to sales
A product team uses AI to summarise hundreds of customer interviews. The summary looks convincing, but it misses what sales hears on calls every week. A named boundary spanner, one product manager who meets sales every two weeks, checks the AI themes against what happens in the field. Because the correction goes back into the prompt and the coding scheme, the learning stays in the workflow instead of in one person's head.
Example 3: Operations escalates an AI forecast
An AI demand forecast suggests cutting stock before a seasonal peak. The planner has doubts but does not know who can override the model. Once the team agrees on an escalation rule (any forecast that moves more than a set threshold goes to a named owner), people stop both obeying the model blindly and ignoring it after one miss.
| Situation | Weak H2H | Strong H2H | Mechanism at work |
|---|---|---|---|
| AI draft rejected by another function | Re-argued case by case | Shared acceptance standard | Aligned goals across levels |
| AI insight contradicts field knowledge | Ignored or blindly accepted | Checked by a named boundary spanner | Boundary spanning |
| Functions read an output differently | Friction becomes conflict | Differences are surfaced and used | Cognitive diversity |
| Someone spots a wrong output | Stays silent | Flags it; the fix goes back into the workflow | Psychological safety |
| Unclear who can override AI | "It depends" | Written escalation rule and owner | Aligned goals / authority |
AI as tool, not teammate
AI removes friction in collaboration — search, drafting, translation between jargons, summarisation across silos. It does not align goals, build trust, or own a decision. People still do that work. Treating AI as a tool (not a teammate) is what lets the four mechanisms above compound instead of dilute.
Fix H2H first, then scale H2M
Every successful AI rollout we've shipped started with a short H2H reset: aligned goals across levels, explicit cross-functional handoffs, named owners for ambiguous calls, and enough safety to surface what isn't working. Then — and only then — does H2M scale cleanly. That is the order. Reverse it and you buy faster silos.
Continue reading
Human-in-the-Loop vs Human-on-the-Loop: How to Choose the Right AI Oversight Model
Human-in-the-loop (HITL) and human-on-the-loop (HOTL) are not interchangeable labels. This guide compares them across ten dimensions, gives four tests for choosing per decision point, walks through healthcare and enterprise use cases, and shows how to implement the right model in 90 days.
How to Audit AI–Human Collaboration Quality: A 4-Dimension Methodology
Most 'human-AI collaboration' audits measure tool usage. This guide shows the four dimensions — goal alignment, boundary spanning, cognitive diversity, psychological safety — and the metrics and questions a professional audit actually uses.
AI Adoption in Healthcare: Why Most Rollouts Fail — and How to Join the 5% That Win
95% of enterprise GenAI pilots produce no measurable impact. In healthcare the stakes go beyond ROI — to patient safety and EU AI Act compliance. Here's what the winning 5% do differently.
The AI Adoption Paradox: Why Investment Is Soaring While Transformation Stalls
92% of companies are increasing AI investment. Only 1% are 'mature' in deployment. The answer lies in organizational psychology — not technology.