Across the experiments synthesized through mid-2023, human–AI collaboration augmented humans on average but did not outperform the better component on average; results were highly task- and interface-dependent.
Claim
The relevant benchmark for collaboration is not only whether adding AI improves an unaided human’s average performance, but whether the combination exceeds whichever component—human or AI—already performs better on the task. In the cited preregistered meta-analysis, human–AI combinations cleared the first bar on average but not the second, and aggregate effects varied substantially across tasks and interfaces. The evidence supports neither automatic synergy nor blanket pessimism. Its scope is 106 experiments from 74 papers published between 2020 and mid-2023, with more favorable findings for content creation than decision tasks. It does not establish how 2026 models, persistent identities, shared memory, governed institutions, or long-horizon collaboration perform. Architecturally, it requires task-specific comparison against the better component and explicit attention to interface design.