AI Tutoring After Mistakes: A Yearlong Randomized Evaluation of NUMI in Middle-School Response to Intervention

Last registered on August 10, 2026

Pre-Trial

Trial Information

General Information

Title
AI Tutoring After Mistakes: A Yearlong Randomized Evaluation of NUMI in Middle-School Response to Intervention
RCT ID
AEARCTR-0019329
Initial registration date
August 07, 2026

Initial registration date is when the trial was registered.

It corresponds to when the registration was submitted to the Registry to be reviewed for publication.

First published
August 10, 2026, 4:58 PM EDT

First published corresponds to when the trial was first made public on the Registry after being reviewed.

Locations

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information

Primary Investigator

Affiliation
University of Toronto

Other Primary Investigator(s)

Additional Trial Information

Status
In development
Start date
2027-08-11
End date
2028-06-30
Secondary IDs
Prior work
This trial does not extend or rely on any prior RCTs.
Abstract
This study is a yearlong, student-level randomized evaluation of NUMI, a mastery-based computer-assisted learning platform for middle-school mathematics, conducted in partnership with Hamilton County Schools (Tennessee) within the district's Response to Intervention (RTI) mathematics program during the 2027-28 school year. The evaluation isolates the marginal value of conversational AI tutoring delivered at the moment a student makes a mistake, holding all other platform features fixed. In both study arms, students practice identical content under identical mastery rules; when a student first submits an incorrect answer, the platform responds with structured help. In the treatment arm, a guard-railed AI tutor guides the student through the error with questions and step-specific support; in the control arm, the platform displays a worked solution. AI support is available only after a mistake and at no other time. The primary questions are whether AI tutoring improves recovery from mistakes, improves delayed retention of practiced material measured on unassisted in-platform 'improvement checks' three to five school weeks after practice, and raises term-level mathematics achievement on the district's NWEA MAP assessment. The study builds on a prior two-year Khanmigo cluster randomized trial and a NUMI pilot in the same district. This pre-analysis plan is registered before randomization.
External Link(s)

Registration Citation

Citation
Oreopoulos, Philip. 2026. "AI Tutoring After Mistakes: A Yearlong Randomized Evaluation of NUMI in Middle-School Response to Intervention." AEA RCT Registry. August 10. https://doi.org/10.1257/rct.19329-1.0
Experimental Details

Interventions

Intervention(s)
All students in both arms use the NUMI platform under identical mastery-progression rules: for each weekly exercise, a student must reach three consecutive correct answers before advancing. The single experimental contrast is the form of structured support delivered after a mistake, where a mistake is defined as a first submitted incorrect (non-skip) answer to an exercise question. Upon a mistake, treatment-arm students receive the NUMI AI tutor - a guard-railed large-language-model tutor that initiates a structured walkthrough of the error, withholds direct answers, and responds to the student in dialogue - while control-arm students receive a worked solution for the question. The trigger (first incorrect submission), the timing (immediately following the mistake), and the surrounding platform experience are identical across arms. The AI tutor is not available at any other point; there is no pre-attempt help channel in either arm. Students work on NUMI approximately one to two sessions per week within their existing RTI mathematics block.
Intervention Start Date
2027-10-18
Intervention End Date
2028-04-07

Primary Outcomes

Primary Outcomes (end points)
Three primary outcome families are pre-specified:

1) Post-mistake recovery, measured at the mistake-event level during delivery weeks: (a) first-attempt correctness on the next new question following the mistake; (b) number of submitted attempts to the next correct answer; (c) school-clock minutes to the next correct answer (capped at 60); and (d) reaching the three-in-a-row mastery threshold on the week's exercise.

2) Delayed retention on NUMI improvement checks: correctness on in-platform, unassisted delayed-assessment items covering weekly exercises practiced during Window 1, stacked across assessment Waves 1 and 2 at the student-by-item level. This is the single pre-registered confirmatory endpoint of the study.

3) MAP mathematics achievement: end-of-window NWEA MAP RIT score (ANCOVA conditioning on beginning-of-window RIT), in population-standard-deviation units, for Window 1 and for the fall-to-spring school year.
Primary Outcomes (explanation)
Improvement checks are in-platform, unassisted assessments administered in four waves; each wave's instrument comprises parallel-form items drawn from the NUMI item bank corresponding to specific prior weekly exercises (practiced item types, not previously seen items), plus parallel-form re-probes of earlier-wave material. Assessment conditions are identical across arms (no AI, no solution display until the full set is submitted). The confirmatory estimand is the intent-to-treat effect of Window-1 AI assignment on improvement-check correctness, estimated by OLS on the stacked student-by-item file with item, wave, and section fixed effects and standard errors clustered at the student level (the unit of randomization); there is a single confirmatory hypothesis and no multiplicity adjustment applies to it. Post-mistake outcomes are estimated at the mistake-event level with weekly-exercise fixed effects, accompanied by per-student aggregates, and summarized by a single standardized recovery index (Anderson 2008). MAP outcomes follow an ANCOVA specification (endline RIT on assignment, baseline RIT, and demographics) in population-SD units, clustering by student. The lag between practice and assessment is defined in school days, targeting a 3-5 school-week lag. Power is calibrated to the 2026 NUMI pilot rather than to assumed design parameters.

Secondary Outcomes

Secondary Outcomes (end points)
Retention trajectory: probe-level correctness as a function of school-day lag and its interaction with treatment, using the multi-lag structure created by re-probed material. Practice-process outcomes: weekly exercises mastered, questions attempted and completed, time on NUMI, and Tuesday-continuation rates (mastery not reached by the end of the Monday session). Window-2 estimates: under the crossover scenario (Scenario A), the opposite-history AI effect on the Wave 3-4 analogue of the confirmatory outcome, the Window-1 by Window-2 sequence interaction, and a pooled within-student crossover estimate reported conditional on the interaction test; under the voice-module scenario (Scenario B), the effect of voice-enabled versus text-only AI tutoring on the same outcome families.
Secondary Outcomes (explanation)
The Window-2 design is governed by a pre-specified readiness rule: if a voice-interaction version of the tutor passes the readiness gate by January 15, 2028, Window 2 randomizes students 50/50 to voice-enabled versus text-only AI (Scenario B); otherwise Window 2 is a balanced within-student crossover of the Window-1 assignment (Scenario A). The scenario decision is documented by registry amendment before Window-2 rostering and is made without reference to any Window-1 outcome data. Secondary outcomes receive Romano-Wolf stepdown p-values alongside unadjusted inference; the post-mistake outcomes are additionally summarized by a single standardized recovery index. Pre-specified exploratory analyses (heterogeneity by baseline MAP, grade, gender, race/ethnicity, economically disadvantaged status, special-education status, and run-in recovery-rate tercile; tutor-engagement measures in the treated arm; and a Khan Academy substitution check) are labeled as exploratory and reported without confirmatory claims.

Experimental Design

Experimental Design
The study is a student-level randomized controlled trial embedded in Hamilton County Schools' RTI mathematics program across approximately 20 middle schools (grades 6-8), with an expected analysis sample of approximately 2,000 students. Each student is randomized once, before Window 1, with half assigned to AI post-mistake tutoring and half to worked solutions. The study year comprises a run-in phase (all students in solutions-only mode; the tutor is frozen and evaluated against a readiness gate) followed by two randomized delivery windows keyed to the district's trimester MAP schedule. Window 1 (approximately eight protected delivery weeks, October-December 2027) is the confirmatory comparison. Window 2 (approximately eight weeks, February-April 2028) follows one of two pre-registered scenarios - a balanced within-student crossover, or a voice-versus-text randomization - selected by a pre-specified readiness rule before Window-2 rostering. Delayed 'improvement checks' are administered in four waves, two per window, at a targeted 3-5 school-week lag. Under either Window-2 scenario, every student receives access to the AI tutor at some point during the year. Teachers are not informed of individual assignments; the tutor introduces itself in-product at each treated student's first post-mistake encounter.
Experimental Design Details
Not available
Randomization Method
Randomization is performed automatically by the NUMI platform's built-in randomization routine at the moment each student account is created, assigning the student 50/50 to the AI-versus-solutions condition. Assignment is unstratified; section identifiers, run-in platform behavior, and baseline MAP enter the analysis as precision covariates rather than as design strata. Accounts are created as students join the RTI roster (from August 2027), so each student's assignment exists from account creation but is inert during the run-in, when both arms receive solutions-only support; the assignment flag activates on October 18, 2027. Mid-year entrants are randomized automatically at entry. The assignment routine (code and version), the full assignment roster with account-creation timestamps, and activation logs will be archived and posted with the registry entry at the close of the study.
Randomization Unit
Individual student. Student-level randomization was chosen over section-level assignment on power grounds: simulation-based calculations using the pilot's variance structure indicate that student-level assignment is required to attain 80 percent power for the pilot-calibrated target effect, whereas section-level randomization of the same design yields power below 65 percent. Within-classroom contamination channels are addressed by teacher blinding, private individualized delivery, and pre-specified contamination checks.
Was the treatment clustered?
No

Experiment Characteristics

Sample size: planned number of clusters
Not applicable - randomization is at the individual student level, so there are no clusters. For context, students are drawn from approximately 20 middle schools (grades 6-8).
Sample size: planned number of observations
Approximately 2,000 students on the RTI mathematics roster across the school year (analysis sample). For the confirmatory outcome, approximately 24 protocol-eligible improvement-check items per student across Waves 1-2, yielding on the order of 48,000 student-by-item observations. Post-mistake analyses are expected to draw on more than 50,000 mistake events per window.
Sample size (or number of clusters) by treatment arms
Window 1: 50/50 individual assignment - approximately 1,000 students assigned to AI post-mistake tutoring and approximately 1,000 students assigned to worked solutions. In Window 2, every student receives AI support: under the crossover scenario each student's Window-2 assignment is the complement of their Window-1 assignment; under the voice scenario students are randomized 50/50 to voice-enabled versus text-only AI tutoring.
Minimum detectable effect size for main outcomes (accounting for sample design and clustering)
Power is calibrated to the 2026 NUMI pilot rather than to assumed parameters (single delayed-item correctness mean 0.37, SD 0.483; within-student cross-item correlation approximately 0.40; target effect 3.2 percentage points). For the confirmatory delayed-item outcome with N approximately 2,000 students and run-in conditioning absorbing 30-50 percent of between-student variance, the minimum detectable effect (80 percent power, 5 percent two-sided) is approximately 2.5-3.0 percentage points, implying power of roughly 82-95 percent against the 3.2 pp target. The MAP ANCOVA (N approximately 2,000, baseline R-squared approximately 0.70) yields an MDE of approximately 0.07 population SD per window. Post-mistake next-question correctness, with event samples above 50,000 per window, yields an MDE below 1.5 percentage points. A simulation-based power appendix using pilot microdata will be posted with the registration.
IRB

Institutional Review Boards (IRBs)

IRB Name
IRB Approval Date
IRB Approval Number
Analysis Plan

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information