Experimental Design
Population and stratification.
The partner firm registered its employees for the study through a recruitment grid crossing role type with tenure. Role types follow the firm's records: Experts (employees on a leadership and technical-authority track), Developers (employees hired to deliver code-based assignments), and Non-coders (employees in adjacent functions such as sales and marketing, typically without formal coding skills). Tenure has three levels: junior (under 3 years), mid (3 to 8 years), and senior (8 years or more). The randomised population is the 6,806 registrants who completed at least one pre-experimental assessment.
Treatments.
Two treatments are randomised independently. AI access is a between-participant factor with three arms (No AI, Weaker AI, Stronger AI, described in the Intervention field). Requirement format is a within-participant crossover: each participant completes one block with formal assignments and one block with stakeholder-interview transcripts, and the order is randomised with probability one half. Experts and Developers are assigned to No AI with probability 1/5 and to each AI arm with probability 2/5. Non-coders are assigned only to the two AI arms with probability 1/2 each, because asking participants without coding skills to complete programming tasks without AI would be unfair and demotivating; their No AI performance is set to zero. All participants receive the same two blocks in the same order (the block content is fixed, the format is not), the seven tasks within a block appear in a randomly shuffled order per participant, and participants choose freely which task to attempt and when.
Tasks.
The 14 experimental tasks were selected from 20 candidates developed with the firm's engineering team, each with unit tests and a scoring rubric validated by the firm. Tasks that the strongest available model solved at the first attempt from the pasted formal assignment were removed, and the remaining tasks were split into two blocks balanced on a machine difficulty index (from an LLM-alone baseline) and a human difficulty index (from expert rankings).
Pre-experimental measurement.
Before the experiment day, every participant completes a battery of assessments: a registration survey, a professional-experience survey, a coding-skills survey, a timed 15-item coding assessment, an LLM-expertise survey, an LLM-driven career interview, and a cognitive and personality battery (Raven's Progressive Matrices, Digit Span, Cognitive Flexibility Inventory, Ten-Item Personality Inventory, General Risk Propensity Scale, and a delay-discounting task). Six confirmatory expertise measures come from this battery: coding knowledge (the timed coding assessment score), career experience (total professional experience in years), on-the-job coding time (the share of the working week spent on the work the tasks exercise), AI expertise (a five-component index of GenAI knowledge and use), fluid intelligence (Raven's score), and working memory (total Digit Span).
Hypotheses.
We test nine confirmatory hypotheses.
- H1: the effect of AI on task performance is positive, and larger for Stronger than for Weaker AI.
- H2: the effect of AI is smaller when requirements arrive as a stakeholder interview rather than a formal assignment, and more so for Stronger AI.
- H3 to H8: the effect of AI is greater when coding knowledge (H3), career experience (H4), on-the-job coding time (H5), AI expertise (H6), fluid intelligence (H7), or working memory (H8) is higher, in each case more so for Stronger AI.
- H9: AI access transforms the relative value of knowledge and cognitive traits, so the effect of AI grows more with an index of cognitive traits (fluid intelligence and working memory) than with an index of knowledge (coding knowledge, on-the-job coding time, and career experience).
Estimation.
The hypotheses are tested in one ordinary-least-squares regression per block. The block score is regressed on the two AI-arm indicators, the format indicator, the six z-scored expertise measures, and the full set of their interactions. H9 uses a companion regression per block that replaces the six measures with three range-normalised indexes (cognitive traits, knowledge, AI expertise). An LLM-alone baseline, in which the same models attempt every task under both formats with three prompt scripts and three repetitions, provides each task's machine difficulty index and the benchmark for value added.
Exploratory analysis.
Heterogeneous effects of both treatments are estimated with a post-selection Lasso (with selective inference) and an honest causal forest over the pre-treatment covariates, with best-linear-projection, group-average-effect, and RATE summaries. Primary analyses include every participant who was assigned a condition; pre-specified robustness subsets are participants with at least one valid submission, participants meeting the engagement threshold, and each participant's first block only.