Big Idea Think TankBring us the question.
Menu

Field paper / Publication decision: publish

Before the model, map the decision

A field paper on what should justify generative AI in consequential knowledge work.

Editorial accountability
Big Idea Think Tank Editorial Function
Published
Modified
Next review

01 / Question

Abstract and question

An institution can buy a capable model and still build a weak decision process around it. This paper asks a narrower question: what evidence should justify moving generative AI from a bounded experiment into consequential knowledge work?

02 / Thesis

Thesis and revision conditions

For consequential knowledge work, a prospective trial of the full human-AI workflow will predict deployment value better than model benchmarks or user-reported speed alone. Deployment value includes decision quality, completion time, the frequency and recovery cost of material errors, human review, and rollback. Revise or reject this thesis if replicated trials in comparable workflows show that benchmarks and self-reports predict those outcomes just as well, or that review and recovery design do not change them.

03 / Evidence

Evidence

  1. Finding 01Source 2

    A preregistered randomized experiment involving 758 Boston Consulting Group consultants found a sharp task boundary. With GPT-4, participants completed 12.2 percent more of 18 within-frontier tasks and finished 25.1 percent faster on average. On one outside-frontier managerial task, AI-assisted participants were 19 percent less likely to reach the correct answer.

    Boundary
    One consulting experiment using GPT-4 as available in June 2023; it does not establish effects for other tasks, populations, or models.
  2. Finding 02Source 3

    A preregistered meta-analysis synthesized 370 effect sizes from 106 human-participant experiments. Human-AI combinations performed better than humans alone on average, but worse than the stronger of the human or AI working alone. Decision tasks showed losses, while creation tasks produced comparatively better results. Task type and baseline performance materially changed the result.

    Boundary
    The review covered studies published from January 2020 through June 2023 and reported substantial heterogeneity and possible publication bias.
  3. Finding 03Source 4

    NIST's Generative AI Profile states that laboratory and benchmark tests may not extrapolate to real-world use. It recommends empirically validated capability claims, measurements under conditions similar to deployment, documented human-oversight roles, ongoing monitoring, source verification, and mechanisms to disengage or deactivate systems whose outcomes conflict with intended use.

    Boundary
    Voluntary, non-sector-specific risk guidance. It does not determine legal compliance or approve a deployment.
  4. Finding 04Source 5

    OMB Memorandum M-25-21 applies to United States executive-branch agencies. For high-impact federal AI uses, it requires pre-deployment testing and impact assessment, periodic human review, monitoring for context changes and adverse effects, suitable human oversight, fail-safes where practicable, and access to human review or appeal when appropriate.

    Primary source
    OMB Memorandum M-25-21
    Boundary
    Federal agency guidance, not a general private-sector legal standard.

04 / Counterevidence

Counterevidence

  1. Finding 01Source 6

    A staggered workplace rollout covering 5,172 customer-support agents increased issues resolved per hour by 15 percent on average. Less experienced and lower-skilled agents improved both speed and quality, while the most experienced agents recorded small speed gains and small quality declines. A stable workflow with repeated problems can produce broad gains, so a blanket presumption against deployment would ignore useful evidence.

    Boundary
    One customer-support setting using an earlier model and a specific assistance design; effects varied by worker skill.
  2. Finding 02Source 7

    In an early-2025 randomized trial, 16 experienced open-source developers completed 246 issues in repositories they knew well. They took 19 percent longer when AI was allowed, yet afterward believed AI had made them 20 percent faster. The authors describe the result as a snapshot of one setting rather than an estimate for software work generally.

    Boundary
    Small, selected sample using early-2025 tools in familiar repositories; it does not generalize to all developers, tools, or tasks.
  3. Finding 03Source 8

    METR's February 2026 follow-up could not estimate the current productivity effect reliably because developers selected out of AI-disallowed work and concurrent agents obscured time measurement. The researchers believed speed gains had probably increased, but described their evidence about the size of that change as very weak. Local trials can expose mistaken beliefs, yet their answers decay as tools and work patterns change.

    Boundary
    The follow-up is a methodological update with unreliable central estimates, not affirmative evidence of a specific current speedup.

05 / Method

Method

We treated the full decision process as the unit of analysis. The source set includes two workplace experiments, one meta-analysis, one developer trial and its methodological update, and two public-governance documents. We separated observed outcomes from official guidance and kept each result inside its studied setting. A practical trial should predeclare the task, affected people, outcome measure, unacceptable error, human-only baseline, AI-only condition where safe, proposed human-AI workflow, review authority, appeal path, stop trigger, and rollback owner. It should measure quality and recovery cost alongside time.

06 / Limits

Limitations and unresolved questions

This paper proposes a decision protocol. It does not report a Big Idea Think Tank engagement or prove a universal effect. The evidence does not cover every sector, current model, population, or consequence. Legal duties, labor agreements, privacy rules, procurement terms, and sector safety requirements sit outside its scope. Model updates, task drift, and changes in who bears an error can invalidate a prior result.

Sources

  1. Launch decisions and editorial recordInternal publication record
  2. Dell'Acqua et al., Navigating the Jagged Technological FrontierOpen primary source
  3. Vaccaro, Almaatouq, and Malone, When combinations of humans and AI are usefulOpen primary source
  4. NIST AI 600-1, Generative Artificial Intelligence ProfileOpen primary source
  5. OMB Memorandum M-25-21Open primary source
  6. Brynjolfsson, Li, and Raymond, Generative AI at WorkOpen primary source
  7. METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer ProductivityOpen primary source
  8. METR, We are Changing our Developer Productivity Experiment DesignOpen primary source

Revision history

  1. Initial publication with primary sources, counterevidence, limitations, a review date, a decision rule, and a commercial disclosure.

07 / Decision

publish

Publish as a testable operating position. This paper does not approve any deployment or claim prior results. Use it to design bounded trials with predeclared stop conditions. Review it by October 30, 2026, and earlier after a material model change, a change in the cited guidance, or the first completed application. Revise if new field evidence changes the predicted value of workflow trials. Close or route the paper if the question becomes a sector-specific legal or safety determination.