Experimental Design
Structure. The trial randomizes at two levels within a single evaluation session.
Across students. We randomize students into one of three arms: pure control, spontaneous use, and encouraged use.
At the question level. We randomize the cost of adopting AI across post-test questions. The questions that can be rendered in either format are split into two sets of approximately equal size. Both sets are intercalated through the section so that consecutive questions come from different sets. Each student is randomly assigned one of the two sets to carry the spoken fragment, while the other set appears fully written. Therefore, participants alternate between written and spoken questions (e.g., text, audio, text, audio). Further, every question appears in the high-cost format for half of the students and in the low-cost format for the other half. This design makes format orthogonal to question content by construction.
A secondary draw establishes the order of the blocks for the three different subjects tested: mathematics, reading, or problem solving, taking one of three values. This order carries across both the pre- and post-tests. This makes copying harder, as described below.
Crossing both draws yields six versions of the evaluation. Then, combining these six test versions with the three different arms results in 18 different cells. Participants occupy exactly one of these cells. These cells depend entirely on the version of Section B and the session arm because, apart from block order, Sections A and C have a fixed format for all students. Figure 1 depicts the experimental design graphically for intuition purposes.
Experimental design
Notes: Within each pair (V1 and V2, V3 and V4, V5 and V6), the same questions change format, and this variation identifies the cost of adopting AI. Across pairs, the block order changes, which only makes copying harder. Sections A and C are the same in every version (except for the block order in Section A). Each square is a question, and shaded squares are spoken; blocks hold more questions than the four drawn. Questions that cannot be rendered as audio are written in every version.
The pre-test as a two-format baseline. Section A contains both written and audio questions, the same for every participant. Our analysis uses the overall pre-test score as the measure of baseline ability. As a secondary analysis, we use the written and the spoken subscores separately.
Field plan. The visitation plan includes 18 Cecyteg campuses, visited once each, six on each of three of the four Saturdays between October 17 and November 7, 2026. Every campus runs three sessions in one day (Saturday), with each session corresponding to one arm. The three sessions run back-to-back. This impedes students from different sessions from having time to interact. A pilot runs on October 3, 2026 at three campuses outside the 18. We pool the pilot data with the main sample only if the instrument, the protocol, and the assignment procedure remain unchanged afterward. If anything changes, we exclude the pilot. With the pilot, the sample would have 21 campuses, 63 sessions, and about 1,575 students.
Copying. We arrange four features in the protocol to limit copying, while avoiding any effect on our identification. First, proctors assign seats to spread students as far apart as the laboratory layout allows on each campus. Second, cardboard partitions between adjacent test stations prevent participants from seeing neighboring monitors. Third, a neighbor participant who draws another block order works on different content. Fourth, the answer options in multiple-choice questions appear in a different order for each student.
Enrollment. Participants sign up voluntarily at their campus before a fixed deadline. After we gather the roster, we randomly assign participants to a session (arm). We also randomly allocate surplus sign-ups, if any, to a session, where they join that session’s waitlist in a random order.
Primary analysis sample. The primary sample includes all students who sit the evaluation in all sessions. Pure-control participants who use AI tools despite the instructions and videos are also kept in the primary sample because dropping these observations would bias random assignment. Finally, we treat recording loss as missing data and report loss rates by arm, as well as no-shows by campus and session (arm).
Estimating Equations. We estimate the average effects of AI access and endorsement on use and performance (intention-to-treat), as well as the effects on performance for those who use AI (treatment-on-the-treated).
Inference. Since students are randomized to arms individually within campus×grade level×gender, each working individually on a partitioned computer and watching the intervention video alone with headphones, we report robust standard errors for the student-level estimates and student-level clustered standard errors for the question-level estimates. We use these as the primary basis for inference and for the q-values. As a robustness check, we also report standard errors clustered at the stratum level (campus by grade level by gender), session level, and campus level. The campus-level clustered standard errors will be estimated using a wild cluster bootstrap (due to the relatively low number of campuses). Finally, we report randomization inference that permutes students’ assignment to sessions using the same randomization algorithm.
Multiple testing. We fix four primary families, defined by outcome (AI use or performance) and by question type (average effects or heterogeneity by ability). First, the AI use family contains β_1 and β_2-β_1 from equation (1), with the share of questions with AI use as the outcome, and the cost effect under spontaneous use relative to pure control θ_4 and its change with the demonstration θ_5-θ_4 from equation (2), with question-level use as the outcome. Second, the performance family contains the same four parameters with the score as the outcome. Third, the heterogeneity in AI use family contains the access-by-ability interaction under spontaneous use δ_3 and its change with the demonstration δ_4-δ_3 from equation (3), with the share of questions with AI use as the outcome, and the cost-by-ability interaction under spontaneous use η_4 and its change with the demonstration η_5-η_4 from equation (5), with question-level use as the outcome. Fourth, the heterogeneity in performance family contains the same four parameters with the score as the outcome. Within each family, we report the sharpened q-values of Anderson (2008). Everything else in this registration is secondary: all other coefficients and contrasts in equations (1) to (5), including β_2, θ_1, θ_2, θ_3, θ_5, 1⁄2 (θ_4+θ_5 ), and the remaining δ, κ, and η terms; any use; the quartile estimates; domain scores; the treatment-on-the-treated estimates; and all mechanism measures. We report any analysis not specified here as exploratory.