Endogenous AI Use and the Labor Productivity Distribution

Last registered on October 07, 2026

Pre-Trial

Trial Information

General Information

Title
Endogenous AI Use and the Labor Productivity Distribution
RCT ID
AEARCTR-0018293
Initial registration date
October 02, 2026

Initial registration date is when the trial was registered.

It corresponds to when the registration was submitted to the Registry to be reviewed for publication.

First published
October 07, 2026, 10:44 AM EDT

First published corresponds to when the trial was first made public on the Registry after being reviewed.

Locations

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information

Primary Investigator

Affiliation
Inter American Development Bank

Other Primary Investigator(s)

Additional Trial Information

Status
In development
Start date
2026-10-03
End date
2026-11-28
Secondary IDs
Prior work
This trial does not extend or rely on any prior RCTs.
Abstract
The existing evidence generally finds that when artificial intelligence (AI) tools are provided to workers, the least skilled gain the most (Brynjolfsson, Li, and Raymond 2025; Noy and Zhang 2023; Peng et al. 2023; Dell’Acqua et al. 2026; Cruces et al. 2026). Of course, that only happens if the least skilled workers adopt these novel tools. However, it is precisely this subset of workers who appear to adopt AI tools the least. This randomized controlled trial measures the effect of endogenous AI adoption on the productivity distribution among 1,350 students at Cecyteg (Colegio de Estudios Científicos y Tecnológicos del Estado de Guanajuato), a network of technical-vocational upper-secondary schools distributed across Guanajuato, the sixth most populous state in Mexico. Eligible human subjects are Cecyteg students enrolled in a three-year program who voluntarily sign up and provide assent, with parental consent where required.
Each participant takes a 135-minute screen-recorded evaluation in a school computer laboratory. The evaluation consists of an unaided pre-test, a short video, a post-test, and a survey on the mechanisms affecting AI use. Participants are paid according to their performance on the pre- and post-tests. Both tests measure skills relevant to the local labor market (i.e., reading, mathematics, and problem solving). Because the screen is recorded and every question is time-stamped, use and performance are observed for each individual task.
We randomize at the student level within campus into three arms: (1) pure control with no tools available (other than a calculator); (2) spontaneous use in which AI is permitted and neither recommended nor discouraged; and (3) encouraged use in which the same permission comes with a video demonstration of an AI tool answering a question similar to those on the test. Further, we randomize the cost of adopting AI at the question level. Namely, a post-test question either appears in full on screen, so submitting it to an AI tool requires only copying and pasting, or has one fragment removed from the screen and played as audio, so a prompt requires listening and restating the question.
The primary outcomes are task-level use, the share of questions on which a student uses AI, and the post-test score. We test whether use declines more steeply with cost for lower-ability students. We also test whether the demonstration narrows or widens the resulting gap. We visit 18 campuses once each, six on each of three of the four Saturdays between October 17 and November 7, 2026. Every campus runs three sessions in one day (one per arm), for a total of 54 sessions (18 per arm) with 20 to 30 students each.
External Link(s)

Registration Citation

Citation
Talamas Marcos, Miguel. 2026. "Endogenous AI Use and the Labor Productivity Distribution." AEA RCT Registry. October 07. https://doi.org/10.1257/rct.18293-1.0
Experimental Details

Interventions

Intervention(s)
The target group consists of students of Cecyteg’s three-year upper-secondary program. All participants take a screen-recorded, incentivized 135-minute evaluation in a school computer laboratory at their respective campus. Students must leave their belongings away from workstations, so they have no access to outside tools (including phones) inside the laboratories. Then, the evaluation follows a fixed order and duration: (1) intake and seating last ≈ 5 minutes; (2) a briefing, consisting of an introductory video and on-screen instructions, lasts 5 minutes; (3) Section A, an unaided pre-test, lasts 45 minutes; (4) an instruction video lasts 2 to 2.5 minutes; (5) Section B, the post-test, lasts 45 minutes; (6) Section C, a survey on mechanisms affecting AI use, lasts 25 minutes; and (7) closing and handover last 6 minutes.
Each section uses the same timer across arms. The clock begins only after the introductory video and on-screen instructions end, and students are told so. Pencil, paper, and a calculator app are always available for all arms. In the pre-test and in pure control, the calculator is the only application a student may use. The access arms allow the use of AI and non-AI (like Microsoft Office) tools.
Students do not need to use the full time for each section, and they can move freely to the next, although they may not go back to a previous section. Within a section, they may return to earlier questions. The protocol includes no scheduled breaks across sections, as students are rested because they take the evaluation on a non-school day (e.g., Saturday).
Sections A and B each cover content on reading, mathematics, and problem solving. Conversely, Section C surveys mechanisms affecting AI use. After this last section, a closing screen shows the total score for Sections A and B, and a staff member pays the corresponding incentive based on that score.
Students receive MXN$150 for attending plus a performance payment based on their share of correct answers across the whole evaluation: MXN$50 above 20%, MXN$100 above 40%, MXN$150 above 60%, and MXN$200 above 80%. These payments do not accumulate. Hence, the performance payment is capped at MXN$200 and total earnings at MXN$350. Because payment covers the whole evaluation, students have an incentive to exert effort on the pre-test (Section A), the dimension along which we examine heterogeneity throughout.
A data agreement between Universidad Anáhuac México and Cecyteg governs the screen recordings and all other data, including the storage and deletion protocol.
The first intervention randomizes access to AI and its endorsement. The final screen before the post-test is a 2- to 2.5-minute video stating the tool rule for the section that follows, Section B in this case. Each student watches the video with individual headphones provided by proctors. The three versions of the instructions video are as follows:
Pure control. The video restates that the post-test follows the pre-test rules. Namely, students may use only pencil and paper, with the computer calculator as the only allowed application.
Spontaneous use. The video communicates to the student that any tool on the computer, including AI assistants, in the post-test. AI is listed alongside other tools such as browsers and the internet, as well as the pencil, paper, and the computer calculator.
Encouraged use. The video grants the same permission and spends the remaining time on a narrated demonstration in which a student solves a problem similar to those on the evaluation in Sections A and B. The demonstration conveys one message: AI answers questions of this kind correctly. The video also shows two uses. First, AI reviews the student’s own answer and provides the correct answer. Second, AI answers the question directly, also correctly. This arm also provides links to four AI assistants and recommends opening them before the post-test starts.
We generate all three videos using AI, with the same synthetic presenter and plain background. Hence, the voice, setting, and delivery are identical across all arms. The video durations are also very similar, so the arms differ only in the rule and the encouragement or demonstration. The two videos without encouragement fill the time gap by restating the same procedural instructions as the introductory video before the pre-test.
The second intervention varies the cost of adopting AI for a task through the format of the questions. We provide two types of questions. A low-cost question appears in full, so a participant can simply copy and paste it into the prompt box of an AI tool. Conversely, a high-cost question has identical content, but part of the question appears as audio instead of written on the screen; thus, the student must listen to it and restate it in the prompt box, if using any AI tools. The two versions of the question vary only in delivery, as the speaker records the audio from the transcript that the written version displays word by word. Participants may replay the audio as they wish. This second intervention alternates question types throughout Sections A and B. In Section A, each question follows a fixed format for everyone, while in Section B it is randomized across students.
Among the questions we designed, some cannot be rendered as audio; specifically, those that contain a table, a diagram, or an expression that must be seen. These types of questions always appear as text for all participants. These questions also count toward the post-test score, but their format is not randomized.
Intervention Start Date
2026-10-17
Intervention End Date
2026-11-07

Primary Outcomes

Primary Outcomes (end points)
1. Task-level use. This indicator captures whether a student sends a prompt request using an AI tool while a given question is on screen. It applies to each post-test question, even if it is only used for one part of a question, and to the two access arms in which AI is permitted. In the pure control arm, AI is not permitted, so the indicator measures non-compliance and we expect it to be close to zero.
2. Share of post-test questions with AI use. The average of task-level use at the student level.
3. Post-test score. This measure is the score on each post-test question and the student’s overall post-test score, built as described below.
Primary Outcomes (explanation)
Use. We machine-code an indicator for AI tool use from the screen recordings for all arms, including the pure control arm for compliance purposes.
The person-level aggregates. There are two aggregates. First, the intensive margin: the share of post-test questions answered with AI assistance. Second, the extensive margin, which indicates whether a student uses an AI tool on any post-test question. Both aggregates build on the same machine-coding of usage.
Units. To measure use, we record accessing an AI tool on the screen for any given question. For questions that have several parts, use covers accessing AI for any of the items. To measure performance, we use each item’s score, which carries the same weight in the subscores for each block (i.e., reading, mathematics, and problem solving).
Score. We automate the scoring of multiple choice items. For open-ended questions, we use a rubric fixed in advance, along with AI agents and human graders. We do this latter set of grades only for a random subsample of participants; we report the correlation between AI-agent scores and human-grader scores.
We standardize the item score to the mean and standard deviation of the pure-control, low-cost items. Further, the post-test score is the equal-weighted mean of the standardized reading, mathematics, and problem-solving scores, not a pooled mean across items.
Scoring for payment purposes is different from the score we use for our analysis. The main reason is that payment is settled on the closing screen, prior to grading open-ended questions. To benefit participants, we take open-ended questions as correct for payment purposes only. Students are not informed about this at any moment; in fact, the instructions clearly state that payment increases with the number of correct answers. Payment increases in steps, one for every 20% of correct answers.

Secondary Outcomes

Secondary Outcomes (end points)
1. Domain standardized scores in reading, mathematics, and problem solving.
2. Opening AI tools before starting the post-test.
3. Time on task per question.
4. Mechanisms from Section C, including self-reported mental effort over the post-test, familiarity with AI tools, prior use of AI, and whether AI came to mind during the post-test.
5. Completion and the number of questions reached.
6. Score and use rate on questions built to penalize AI use, such as items in which using AI tools costs more time than it saves.
Secondary Outcomes (explanation)

Experimental Design

Experimental Design
Structure. The trial randomizes at two levels within a single evaluation session.
Across students. We randomize students into one of three arms: pure control, spontaneous use, and encouraged use.
At the question level. We randomize the cost of adopting AI across post-test questions. The questions that can be rendered in either format are split into two sets of approximately equal size. Both sets are intercalated through the section so that consecutive questions come from different sets. Each student is randomly assigned one of the two sets to carry the spoken fragment, while the other set appears fully written. Therefore, participants alternate between written and spoken questions (e.g., text, audio, text, audio). Further, every question appears in the high-cost format for half of the students and in the low-cost format for the other half. This design makes format orthogonal to question content by construction.
A secondary draw establishes the order of the blocks for the three different subjects tested: mathematics, reading, or problem solving, taking one of three values. This order carries across both the pre- and post-tests. This makes copying harder, as described below.
Crossing both draws yields six versions of the evaluation. Then, combining these six test versions with the three different arms results in 18 different cells. Participants occupy exactly one of these cells. These cells depend entirely on the version of Section B and the session arm because, apart from block order, Sections A and C have a fixed format for all students. Figure 1 depicts the experimental design graphically for intuition purposes.

Experimental design
Notes: Within each pair (V1 and V2, V3 and V4, V5 and V6), the same questions change format, and this variation identifies the cost of adopting AI. Across pairs, the block order changes, which only makes copying harder. Sections A and C are the same in every version (except for the block order in Section A). Each square is a question, and shaded squares are spoken; blocks hold more questions than the four drawn. Questions that cannot be rendered as audio are written in every version.
The pre-test as a two-format baseline. Section A contains both written and audio questions, the same for every participant. Our analysis uses the overall pre-test score as the measure of baseline ability. As a secondary analysis, we use the written and the spoken subscores separately.
Field plan. The visitation plan includes 18 Cecyteg campuses, visited once each, six on each of three of the four Saturdays between October 17 and November 7, 2026. Every campus runs three sessions in one day (Saturday), with each session corresponding to one arm. The three sessions run back-to-back. This impedes students from different sessions from having time to interact. A pilot runs on October 3, 2026 at three campuses outside the 18. We pool the pilot data with the main sample only if the instrument, the protocol, and the assignment procedure remain unchanged afterward. If anything changes, we exclude the pilot. With the pilot, the sample would have 21 campuses, 63 sessions, and about 1,575 students.
Copying. We arrange four features in the protocol to limit copying, while avoiding any effect on our identification. First, proctors assign seats to spread students as far apart as the laboratory layout allows on each campus. Second, cardboard partitions between adjacent test stations prevent participants from seeing neighboring monitors. Third, a neighbor participant who draws another block order works on different content. Fourth, the answer options in multiple-choice questions appear in a different order for each student.
Enrollment. Participants sign up voluntarily at their campus before a fixed deadline. After we gather the roster, we randomly assign participants to a session (arm). We also randomly allocate surplus sign-ups, if any, to a session, where they join that session’s waitlist in a random order.
Primary analysis sample. The primary sample includes all students who sit the evaluation in all sessions. Pure-control participants who use AI tools despite the instructions and videos are also kept in the primary sample because dropping these observations would bias random assignment. Finally, we treat recording loss as missing data and report loss rates by arm, as well as no-shows by campus and session (arm).
Estimating Equations. We estimate the average effects of AI access and endorsement on use and performance (intention-to-treat), as well as the effects on performance for those who use AI (treatment-on-the-treated).
Inference. Since students are randomized to arms individually within campus×grade level×gender, each working individually on a partitioned computer and watching the intervention video alone with headphones, we report robust standard errors for the student-level estimates and student-level clustered standard errors for the question-level estimates. We use these as the primary basis for inference and for the q-values. As a robustness check, we also report standard errors clustered at the stratum level (campus by grade level by gender), session level, and campus level. The campus-level clustered standard errors will be estimated using a wild cluster bootstrap (due to the relatively low number of campuses). Finally, we report randomization inference that permutes students’ assignment to sessions using the same randomization algorithm.
Multiple testing. We fix four primary families, defined by outcome (AI use or performance) and by question type (average effects or heterogeneity by ability). First, the AI use family contains β_1 and β_2-β_1 from equation (1), with the share of questions with AI use as the outcome, and the cost effect under spontaneous use relative to pure control θ_4 and its change with the demonstration θ_5-θ_4 from equation (2), with question-level use as the outcome. Second, the performance family contains the same four parameters with the score as the outcome. Third, the heterogeneity in AI use family contains the access-by-ability interaction under spontaneous use δ_3 and its change with the demonstration δ_4-δ_3 from equation (3), with the share of questions with AI use as the outcome, and the cost-by-ability interaction under spontaneous use η_4 and its change with the demonstration η_5-η_4 from equation (5), with question-level use as the outcome. Fourth, the heterogeneity in performance family contains the same four parameters with the score as the outcome. Within each family, we report the sharpened q-values of Anderson (2008). Everything else in this registration is secondary: all other coefficients and contrasts in equations (1) to (5), including β_2, θ_1, θ_2, θ_3, θ_5, 1⁄2 (θ_4+θ_5 ), and the remaining δ, κ, and η terms; any use; the quartile estimates; domain scores; the treatment-on-the-treated estimates; and all mechanism measures. We report any analysis not specified here as exploratory.
Experimental Design Details
Not available
Randomization Method
We randomize days before the evaluation, after receiving the student list, using Stata code with seeds fixed in advance. Within each campus, we randomly assign students to one of the three sessions, stratifying by grade level and gender. Each session is one arm, and we randomly draw the order in which the three arms run during the day for each campus. This prevents confounding between the arm and time of day. Each student is also assigned at random one of the two question-format assignments, balanced within session, together with one of three block orders that serve only to make copying harder.
Randomization Unit
Two randomizations occur, both at the student level. First, we randomly assign each student to one of three arms (pure control, spontaneous use, or encouraged use), within campus, grade level, and gender. Second, each student is randomly assigned one of two question-format versions of the post-test. The questions that can be delivered in either format are divided into two sets; in one version the first set carries the audio fragment and the second is fully written, and in the other version the reverse. The format of a given question therefore varies across students, and each student receives both formats. Hence, about half of students answer each question in the audio format and the other half in the fully written format.
Was the treatment clustered?
No

Experiment Characteristics

Sample size: planned number of clusters
Although the treatment is not clustered, we will also report clustered standard errors at the session level (54) and the campus level (18).
Sample size: planned number of observations
Students are randomized individually to arms within each campus. Since each arm runs as one session per campus, students in the same session share common conditions, but we limit this by design. The intervention is delivered via video through individual headphones; timers and instructions are automated at the individual level; and proctors provide only basic instructions and supervision. However, acknowledge that certain session conditions may affect all exam takers, which we address in the inference.
Sample size (or number of clusters) by treatment arms
1,350 students in 54 sessions of 20 to 30 students each, planned at the midpoint of 25, for a range of 1,080 to 1,620. Each student answers approximately 40 post-test questions, which yields approximately 54,000 observations for the question-level analyses and 1,350 for the student-level analyses. The cost contrast rests on the subset of those questions that can be delivered in either format.
Minimum detectable effect size for main outcomes (accounting for sample design and clustering)
Table 1 reports minimum detectable effects (MDEs) for the primary parameters at α=0.05 (two-sided) and 80% power, with N = 1,350 students split equally across the three arms (450 per arm, at 25 students per session). Score MDEs are in standard deviations of the pure-control post-test score, and use MDEs are in percentage points. The first two columns use robust standard errors and include no design effect, since students are randomized individually. The last two columns use standard errors clustered at the campus level and assume an intraclass correlation of 0.05 at the session level. Minimum detectable effects Robust Campus-clustered Parameter Use Score Use Score β_1 (spontaneous vs. control) 5.4 0.15 8.5 0.23 β_2-β_1 (change with demonstration) 7.5 0.15 11.8 0.23 θ_4 (cost effect, spontaneous vs. control) 4.6 0.15 4.9 0.15 θ_5-θ_4 6.5 0.15 6.9 0.15 δ_3 (access by ability) 5.6 0.15 6.0 0.16 δ_4-δ_3 7.7 0.15 8.2 0.16 η_4 (cost by ability) 4.6 0.14 4.9 0.14 η_5-η_4 6.5 0.14 6.9 0.14
IRB

Institutional Review Boards (IRBs)

IRB Name
Comité de Ética en Investigación, Universidad Anáhuac México
IRB Approval Date
2026-04-21
IRB Approval Number
DINV210426
Analysis Plan

Analysis Plan Documents