Research  /  The Harness Is the Variable: Why the Ten-Fold Disagreement Aโ€ฆ

The Harness Is the Variable: Why the Ten-Fold Disagreement About AI Productivity Is Not About AI

Authors SomaSoft Research (prepared by Claude Code for the AURI project)
Published 2026-08-21
SAGL-1.0 preprint Open Access
View License
๐Ÿ“‹ Cite this paper
SomaSoft Research (prepared by Claude Code for the AURI project). (2026-08-21). "The Harness Is the Variable: Why the Ten-Fold Disagreement About AI Productivity Is Not About AI". SOMAsoft Research. Available at https://somasoft.ai/papers/the-harness-is-the-variable. Licensed under SAGL-1.0.

The Harness Is the Variable

Seven randomized trials, one model, and why the ten-fold disagreement about AI's economic effect is not a disagreement about AI

Evidence gate: truthiness 1.000 ยท 10/10 load-bearing claims grounded ยท 14 sources

Best measured gain โˆ’56% development time (Copilot)
Worst measured โˆ’19% expert developers on their own repo
Perception gap 39 points โ€” felt +20%, measured โˆ’19%
Macro spread 10ร— โ€” 0.66% vs 7%

01 โ€” The decomposition

A horse is power. A rider is intent. The harness is the thing that turns one into the other, and it is the only part anyone designs. Take that seriously and the confusing productivity literature stops being confusing.

realised gain = (horse ร— harness) ร— steering โˆ’ verification cost

The terms multiply rather than add, which is the whole point. A superb horse in a bad harness does not deliver a reduced benefit; it delivers a negative one, because someone still pays the cost of checking work that never arrives usable. That is not a theoretical worry. Two of the seven trials below measured exactly that.


02 โ€” Seven trials, one model

Trial Horse Harness Rider Model Published
Peng โ€” GitHub Copilot 0.90 0.95 0.70 +42% โˆ’56% dev time
Noy & Zhang โ€” writing (n=453) 0.90 0.85 0.65 +37% โˆ’40% time, +0.45 SD quality
Brynjolfsson โ€” novice agents 0.85 0.85 0.25 +33% +34%
Brynjolfsson โ€” expert agents 0.85 0.85 0.90 +15% ~0%
BCG โ€” inside frontier (n=758) 0.85 0.80 0.75 +22% +25.1% speed, +12.2% tasks
BCG โ€” outside frontier 0.25 0.80 0.30 โˆ’21% โˆ’19% correctness
METR โ€” experts, own repo 0.85 0.30 0.85 โˆ’20% โˆ’19% speed

Factor scores are the author's reading of each study's setup, not published values. The model output is ordinal โ€” fitted to reproduce direction, not magnitude โ€” and the published outcomes are in mixed units.

The only real test of a model like this is whether it gets the signs right without being told. It does: negative exactly where the trials went negative, near-flat for Brynjolfsson's experienced agents, strongly positive where the coupling was tight.

The two failures fail for different reasons, and the decomposition separates them. BCG's negative is a rider failure โ€” the wrong side of the frontier. METR's is a harness failure โ€” a large private codebase whose context could not be transmitted.

Same horse, same rider, opposite outcome

Copilot and METR both used frontier models with capable engineers. Copilot scored roughly 0.90 on capability with a 0.95 harness โ€” the model sits inside the editor, and the context is already there. METR scored 0.85 on capability with a 0.30 harness โ€” a chat window beside a large repository whose context cannot cross the gap.

One produced a 56% reduction in development time. The other made developers 19% slower. The only large difference between those two columns is the harness.

The finding that should worry an executive. METR's developers were 19% slower and believed they were 20% faster โ€” a 39-point gap between felt and measured productivity. This means self-reported AI productivity gains are worthless as evidence, and that organisations will confidently scale deployments that are costing them money. Any business case built on "our teams say it helps" rests on the one number the literature has directly falsified.


03 โ€” The estimate

Acemoglu puts the decade effect at 0.66% TFP; Goldman Sachs at 7% of global GDP. A tenfold disagreement between serious people is usually a sign they are computing different things. They are.

decade TFP gain โ‰ˆ ฯƒ (share of tasks affected) ร— ฮณ (average realised saving)

Back out each side's implied ฮณ:

Forecast ฯƒ assumed Published Implied ฮณ Matches
Acemoglu (NBER w32487) 5% 0.66% TFP 13.2% Brynjolfsson's measured 14% average
Goldman Sachs 25% 7% GDP 28.0% BCG's 25.1% inside-frontier figure

Neither is wrong about its own evidence. Acemoglu is pricing the average case across a narrow slice of work. Goldman is pricing the good case across a wide one. The argument is about ฯƒ and ฮณ โ€” both empirical questions, not matters of opinion about AI.

ฯƒ โ†“ ฮณ โ†’ 10% 14% 20% 25% 30%
5% 0.50% 0.70% 1.00% 1.25% 1.50%
10% 1.00% 1.40% 2.00% 2.50% 3.00%
15% 1.50% 2.10% 3.00% 3.75% 4.50%
20% 2.00% 2.80% 4.00% 5.00% 6.00%
25% 2.50% 3.50% 5.00% 6.25% 7.50%

Decade TFP gain. Acemoglu occupies the top-left; Goldman the bottom-right.

The estimate. Taking ฯƒ at 10โ€“15% and ฮณ at 14โ€“25%, the defensible decade range is 1.4% to 3.8% TFP, centred near 2โ€“3%. Real, worth having, and nothing like a discontinuity.

But the number is less interesting than its structure. ฯƒ is mostly given; ฮณ is built. Moving ฮณ from 14% to 25% at ฯƒ=15% takes the decade gain from 2.10% to 3.75% โ€” 1.8ร— more output with no change in model capability whatsoever. That delta is interfaces, task selection, verification design and training.

The productivity boost is substantially a deployment choice, not a forecast to await.


04 โ€” How the proceeds should be shared

The default proposal is a sovereign wealth fund on the Alaska model โ€” the Permanent Fund has paid residents roughly $1,600 a year since 1976; Norway's fund exceeds $1 trillion and spends about 3% annually. Senator Sanders proposed an AI version in 2026, and sixteen Nobel laureates warned about displacement. The instinct is right. The mechanism, copied directly, does not work.

First, there is no royalty base. Oil funds charge for extraction of something the public already owns. Compute has no equivalent. And the funding problem is concrete: sovereign funds pay out of cash โ€” royalties or realised gains โ€” while the major AI firms largely do not pay dividends and would need retained earnings or surplus free cash flow before they could.

Second, and more important: most of the surplus will not sit in AI companies. It will sit in the hundreds of thousands of ordinary firms that used AI to avoid a hire. Taxing "AI companies" to capture the gains of AI is like taxing tractor manufacturers to capture the gains of mechanised agriculture โ€” the manufacturers booked a rounding error next to the surplus that accrued across every farm and, eventually, every food buyer. A levy aimed at the labs will miss the money almost entirely.

Who the evidence says to help

The distributional fact is already measured and unusually clear: gains were 34โ€“35% for the least experienced workers and roughly zero for the most experienced. AI raises the floor far more than the ceiling. It compresses the skill premium.

That cuts both ways, and the second edge is sharper. If a novice with AI performs like a mid-level worker, the novice becomes more valuable and more substitutable in the same motion. The scarce complement stops being craft skill and becomes rider skill โ€” knowing which side of the frontier you are on. That is the capability whose distribution decides whether the next decade's gains reach people at all.

The proposal

  1. Capture broadly, not narrowly. Tax the surplus where it lands, through existing corporate and income machinery indexed to the labour-share shift โ€” not through a levy on a dozen model providers. Follow the money to the firms that saved it.
  2. Distribute roughly half as cash. A dividend is honest, simple and politically durable โ€” Alaska's has survived fifty years. It is also small: $1,600 a year is dignity, not transformation. Do not oversell it.
  3. Distribute the other half as capability. Rider training, access to a working harness, and local compute a community owns rather than rents. Cash compensates people for a transition; capability lets them participate in it. Only the second changes ฮณ, and ฮณ is what generates the surplus in the first place.
  4. Fund the harness as public infrastructure. If ฮณ is the policy variable, then interfaces, verification tooling and task-selection training are infrastructure in the same sense roads are. Left purely to firms, ฮณ improves where margins are thickest โ€” which is not where the 34% gains were found.

The through-line. The horse is bought, the same one, by everybody. What differs โ€” between two development teams, between Acemoglu and Goldman, between one country and another โ€” is the harness and the rider. That is where the productivity is, and it is therefore where the distributive question actually lives. Sharing the proceeds and creating them turn out to be the same problem.


05 โ€” Limits


Sources

  1. Dell'Acqua, McFowland, Mollick, Lifshitz-Assaf, Kellogg, Rajendran, Krayer, Candelon & Lakhani, Navigating the Jagged Technological Frontier, Organization Science (BCG field experiment, n=758).
  2. Same, outside-frontier condition: 19% less likely to be correct.
  3. Noy & Zhang, Science, 2023 (n=453).
  4. Brynjolfsson, Li & Raymond, customer-support field study (n=5,172).
  5. Peng et al., GitHub Copilot randomized trial.
  6. METR developer productivity trial, 2025.
  7. Acemoglu, The Simple Macroeconomics of AI, NBER working paper w32487, May 2024.
  8. Goldman Sachs global GDP forecast, via the AEI exchange with Acemoglu.
  9. Alaska Permanent Fund (est. 1976); Norges Bank Investment Management fiscal rule.
  10. Fortune, July 2026 (Sanders proposal, 16 Nobel laureates); TechTimes, June 2026.
  11. AURI project, AGI Capability Probe, 24 July 2026.

Claims gated by the project's truthiness filter. Model reproducible via papers/productivity_model/model.py. Analysis, not investment or policy advice.