FIELD NOTE / EVIDENCE REVIEW 0019 OCT 2026 · 5 PRIMARY SOURCES
← EXPERIMENTS & ESSAYS

Does AI coding
actually make us faster?

A review of trial evidence, self-reports and benchmarks, with explicit limits on what each result can establish.

?
READ THIS FIRST

Evidence from a specific study is not a forecast about every developer. We separate measured outcomes, participant beliefs and changes in the tools being tested.

Start by defining faster

A study of software productivity needs an endpoint. Faster to the first plausible code sample is not the same as faster to a reviewed, tested and merged change. AI-assisted programming may alter each phase differently: drafting, searching, debugging, validation, integration and future maintenance. Without defining the outcome, the claim “AI makes developers faster” is too imprecise to evaluate.

In 2025, the Stack Overflow Developer Survey reported that more respondents distrusted than trusted AI-generated answers; respondents also described frustration with near-correct code. These responses document perceptions and reported behavior. They do not directly measure time saved in production.[1]

An experiment that contradicted expectations

METR's July 2025 randomized controlled study compared work on familiar open-source repositories with and without AI assistance. Sixteen experienced developers worked on 246 tasks. In the study setting, allowing early-2025 AI tools was associated with 19% longer task completion time, even though participants expected AI to accelerate their work.[2]

The result matters because time was measured in an actual development setting rather than inferred from a synthetic code-generation benchmark. But the scope is narrow: experienced contributors, their own established projects, the available tools at that moment and the chosen tasks. It does not establish a universal slowdown for all programmers, projects or later models.

What a single effect estimate leaves out

Individual workloads can vary dramatically. Repeated tasks, unfamiliar repositories, beginner projects and tightly scoped transformations may behave differently. A responsible summary reports sample size, task characteristics, uncertainty and study timing alongside the headline estimate.

The 2026 follow-up adds uncertainty, not a simple reversal

In February 2026, METR described problems with a newer experiment using later AI tools. Developers increasingly declined to participate when tasks prohibited AI use, introducing selection effects; parallel-agent workflows also complicated time measurement. METR judged the newer data an unreliable estimate of the productivity effect of contemporary tools.[3]

The update reported some raw evidence of speedup, but cautioned that the true size was uncertain. This is neither confirmation that AI always helps nor justification for ignoring the earlier trial. It demonstrates why rapidly changing technology and nonrandom participation complicate comparisons.

Our editorial position is methodological: both studies belong in the record, and their populations, tools and measurement problems must remain attached to their findings.

Benchmarks measure their instruments too

Benchmarks can capture important capabilities, but a score is meaningful only if the tasks are valid and the evaluation data remain sufficiently independent from model training. In February 2026, OpenAI said that SWE-bench Verified had become unsuitable for evaluating frontier coding systems, citing issues such as tests that rejected valid solutions and exposure of benchmark material.[4]

That statement comes from a model provider and should be evaluated as a source with its own incentives, even when it presents substantive technical analysis. The lesson extends beyond a single benchmark: model scores should be accompanied by dataset versions, leakage analysis, held-out tests and external replication.

Passing tests and delivering reliable software are different claims

Independent review, security properties, integration constraints and maintainability are difficult to compress into a single number. NIST's Secure Software Development Framework provides a wider vocabulary for development security practices, including preparation, protection, producing secure software and responding to vulnerabilities.[5]

What evidence would change the conclusion?

We would give substantial weight to preregistered studies across varied developer experience levels and repositories, with matched tasks, current tools, enough statistical power and transparent reporting of failures. Useful metrics include elapsed time to accepted change, review effort, regressions, resource cost, security findings and maintenance outcomes over time.

Industry surveys can complement, but not replace, measured trials. Product demonstrations can establish possibility, not population-wide impact. Benchmarks can highlight model progress, not settle deployment economics. A credible future picture will require several kinds of evidence pointing in compatible directions.

Until then, the useful answer is conditional: AI programming assistance may help in some contexts and hinder in others; the size and direction of the effect must be measured against a well-defined task and outcome.

Forecast watch: record the conditions before the outcome

This is a research agenda, not a scored forecast. Before futureofcode.site publishes any probability or dated prediction, its forecast register should specify the exact proposition, deadline, eligible evidence, independent resolution process and versioned updates.

We will keep the original statement visible when an outcome arrives. An accurate forecast is not merely a sentence that can be interpreted generously after the fact. A failed forecast remains part of the public record and can be a valuable result.

EDITORIAL / VERSION RECORD

Revision record

Edition 1.0 — 9 October 2026. Initial evidence synthesis, source references and limitations published. Subsequent material corrections will specify the affected claim, evidence and revision date.

Read our corrections protocol ↗
RESEARCH / PROVENANCE

Works Cited

05 VERIFIED SOURCE ENTRIES
  1. 01
    2025 Developer Survey — AI usage and trustStack Overflow · 2025 · Survey; respondents are not a random sample of all software developers.
  2. 02
    Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer ProductivityMETR — Becker, Rush, Barnes and Rein · 2025-07-10 · Randomized trial; 16 experienced contributors, 246 tasks in familiar repositories.
  3. 03
    We Are Changing Our Developer Productivity Experiment DesignMETR — Becker, Rush, Cunningham, Rein and Mahamud · 2026-02-24 · Follow-up methodological note; selection effects and unreliable newer effect estimate.
  4. 04
    Why We No Longer Evaluate SWE-bench VerifiedOpenAI · 2026-02-23 · Provider-authored benchmark audit; note source perspective.
  5. 05
    SP 800-218 — Secure Software Development Framework, Version 1.1NIST · 2022-02-03 · Established security-practice standard, not a study of AI coding productivity.

References document the evidence considered for this edition. Published findings can change, and sources may disagree.

THE RECORD STAYS OPEN

Have conflicting evidence?

Submit a study, challenge an interpretation, or identify an unreported limitation. Evidence is reviewed before any editorial change.

Improve the record ↗