Adaptive Math Practice and Teacher Data Use: A Randomized Evaluation of AI-Enabled Learning in Uzbekistan

Last registered on August 04, 2026

Pre-Trial

Trial Information

General Information

Title
Adaptive Math Practice and Teacher Data Use: A Randomized Evaluation of AI-Enabled Learning in Uzbekistan
RCT ID
AEARCTR-0019200
Initial registration date
July 27, 2026

Initial registration date is when the trial was registered.

It corresponds to when the registration was submitted to the Registry to be reviewed for publication.

First published
August 04, 2026, 8:55 AM EDT

First published corresponds to when the trial was first made public on the Registry after being reviewed.

Locations

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information

Primary Investigator

Affiliation
World Bank

Other Primary Investigator(s)

PI Affiliation
World Bank
PI Affiliation
World Bank
PI Affiliation
World Bank
PI Affiliation
World Bank
PI Affiliation
World Bank

Additional Trial Information

Status
In development
Start date
2026-09-02
End date
2026-12-31
Secondary IDs
Prior work
This trial does not extend or rely on any prior RCTs.
Abstract
This study evaluates an AI-enabled, curriculum-aligned mathematics platform in public schools in Uzbekistan. Schools will be randomly assigned to business-as-usual instruction, to replace one weekly Grade 5 mathematics lesson with adaptive practice on Eduten, or to receive the same platform plus structured support for teachers to use platform-generated information during their other mathematics lessons. The study will estimate effects on independently assessed mathematics achievement and examine whether structured teacher use of formative data produces gains beyond student platform use alone. It will also measure implementation, instructional practices, teachers’ knowledge of student learning, and potential unintended effects on students and teachers.
External Link(s)

Registration Citation

Citation
Babakhodjaeva, Victoriya et al. 2026. "Adaptive Math Practice and Teacher Data Use: A Randomized Evaluation of AI-Enabled Learning in Uzbekistan." AEA RCT Registry. August 04. https://doi.org/10.1257/rct.19200-1.0
Sponsors & Partners

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information
Experimental Details

Interventions

Intervention(s)
The intervention replaces one of five weekly Grade 5 mathematics lessons with a supervised Eduten session during regular school hours. Eduten provides curriculum-aligned adaptive exercises, immediate feedback, and structured practice. Treatment 1 receives the standard implementation model. Treatment 2 receives the same student-facing model plus structured support for teachers to review and use platform-generated information about student progress and learning gaps during their remaining mathematics lessons.
Intervention Start Date
2026-09-02
Intervention End Date
2026-12-18

Primary Outcomes

Primary Outcomes (end points)
Student mathematics achievement at endline, measured by the overall IRT score on an independently administered Grade 5 mathematics assessment.
Primary Outcomes (explanation)
The independent assessment will cover foundational mathematics content from Grades 1–4 and Grade 5 curriculum content expected during the intervention. Item responses will be scored using an item-response-theory model. The overall endline IRT score will be oriented so that higher values indicate stronger achievement and standardized for interpretation in standard-deviation units using the control-group endline distribution. The corresponding baseline IRT score will be used as a precision-improving covariate.

Secondary Outcomes

Secondary Outcomes (end points)
1. Foundational numeracy IRT score (the items from the overall assessment that will contribute to this sub-score will be pre-registered ahead of the endline data collection)
2. Grade 5 curriculum-aligned mathematics IRT score (the items from the overall assessment that will contribute to this sub-score will be pre-registered ahead of the endline data collection)
3. Common Eduten quiz or assessment scores among the broader Grade 5 population in Treatment 1 and Treatment 2 schools.
4. Teacher accuracy in assessing student performance.
5. Observed instructional practices and data-informed instruction.
6. Teacher-reported data use, workload, stress, and professional judgment.
7. Student mathematics motivation, anxiety, frustration, discouragement, embarrassment, perceived labeling, and perceived teacher attention.
8. Implementation and fidelity measures from platform and field data.
Secondary Outcomes (explanation)
The two independent-assessment domain scores will be constructed using prespecified item mappings and IRT scoring. The items from the overall assessment that will contribute to each of these two sub-scores will be pre-registered ahead of the endline data collection. Teacher prediction accuracy will compare class- or student-level predictions made before endline results are available with realized endline performance. Classroom-practice outcomes will use selected TEACH modules and clearly separated study-specific indicators. Related survey items may be combined into standardized indices, with item membership and sign conventions fixed before treatment-effect analysis. Platform outcomes will use a common assessment presented comparably in Treatment 1 and Treatment 2. Because control students will not use Eduten, these platform outcomes will be used for the Treatment 2 versus Treatment 1 contrast only.

Experimental Design

Experimental Design
The study is a three-arm, school-cluster randomized controlled trial in 450 eligible public schools in Tashkent City and Tashkent Region. Schools will be assigned in equal proportions to business-as-usual control, Eduten adaptive mathematics practice, or Eduten plus structured teacher use of platform-generated learning data. One Grade 5 class per school will be randomly selected for the independent baseline and endline assessments, teacher survey, and classroom-level measurement. The main analysis will estimate intention-to-treat effects based on school assignment.
Experimental Design Details
Not available
Randomization Method
Randomization will be performed by computer. Schools will be randomized within prespecified blocks defined by geography and school size and, if available and sufficiently complete, prior school performance. The randomization code, input frame, seed, block definitions, and assignment output will be archived. Within schools with multiple Grade 5 classes, survey software will randomly select one eligible class on site.
Randomization Unit
School. One Grade 5 class per school will subsequently be sampled for intensive independent measurement. Treatment is assigned to the school and applies to the Grade 5 population in treatment schools.
Was the treatment clustered?
Yes

Experiment Characteristics

Sample size: planned number of clusters
450 schools
Sample size: planned number of observations
Approximately 12,150 students in the independent assessment sample, assuming one randomly selected Grade 5 class averaging 27 students in each of 450 schools. Platform-based outcomes are expected to cover approximately 18,000-22,500 Grade 5 students across the 300 treatment schools, depending on realized Grade 5 enrollment and data availability. Approximately one Grade 5 mathematics teacher and one sampled class will contribute teacher- and classroom-level measures per school.
Sample size (or number of clusters) by treatment arms
150 control schools, 150 Treatment 1 schools, and 150 Treatment 2 schools. The independent assessment is expected to include approximately 4,050 students per arm if sampled classes average 27 students. Platform outcomes will be available only in the two treatment arms and may cover approximately 60–75 Grade 5 students per treatment school.
Minimum detectable effect size for main outcomes (accounting for sample design and clustering)
Power calculations assume 150 schools per arm, 80 percent power, a two-sided 5 percent significance level, an intraclass correlation of 0.20, a coefficient of variation in cluster size of 0.50, and equal allocation across three arms. For the independently administered baseline and endline assessments, approximately 27 students will be assessed per school. The corresponding minimum detectable effect size is approximately 0.16 standard deviations without precision gains from baseline data, 0.15 standard deviations if baseline covariates explain 5 percent of outcome variation, and 0.15 standard deviations if they explain 15 percent. For the common platform-based assessments administered to approximately 75 students per school in the two treatment arms, the minimum detectable effect size is approximately 0.15 standard deviations without baseline precision gains, 0.15 standard deviations with 5 percent baseline explanatory power, and 0.14 standard deviations with 15 percent baseline explanatory power. Thus, the study is designed to detect pairwise effects of approximately 0.14–0.16 standard deviations, depending on the outcome sample and the predictive power of baseline data. These calculations apply to unadjusted pairwise comparisons and will be updated using realized cluster sizes, outcome availability, and baseline explanatory power.
IRB

Institutional Review Boards (IRBs)

IRB Name
IRB Approval Date
IRB Approval Number