Experimental Design Details
Our primary research questions focus on how our intervention moves beliefs and behavior through the lens of the theory of change. To do so, we test for treatment effects first on teacher beliefs about child development, then on teacher behavior measured through teacher-child interactions, and finally on child outcomes. Secondary research questions include heterogeneity of treatment effects by classroom and child characteristics, changes to disparities in child development, changes in children's skills measured through other algorithms applied to the audio data, and spillovers to parental knowledge and investment.
A primary question is whether our intervention changes teacher beliefs about child development. Our primary measure of teacher beliefs will be teacher scores on the SPEAK-CAT. SPEAK-CAT is a computer-adaptive test that provides a measure of knowledge of child development in a short administration time. Teachers complete the SPEAK-CAT prior to the start of the intervention (``baseline"), approximately halfway through the study period (``midline"), and at the end of the nine-month study (``endline").
In addition, we will elicit secondary measures of teacher beliefs during teacher surveys at baseline, midline, and endline. These include teacher beliefs about growth mindset, teacher beliefs about AI, and teacher reports about their relationship with the children in class (Student Teacher Relationship Scale - Closeness Factor (Pianta, 2001)).
We investigate whether our intervention changes teacher-child interactions. We measure interactions by applying algorithms to the Luet audio recordings to detect conversational turn counts (CTCs) between an adult and a key child (the child wearing the given Luet). The algorithm uses speaker identification to detect when an adult is speaking and when the key child is speaking, and uses standard counting rules from the literature to calculate CTCs. We test whether our intervention changes language inputs by comparing CTCs in the treatment and control groups in the post-intervention period. Our secondary language input is adult word count (AWC), measured as the number of words spoken by an adult to the key child and calculated based on the audio recordings obtained from Luet as well.
We measure child outcomes during the baseline period (``baseline"), approximately halfway through our study (``midline"), and at the end of the study (``endline"). The primary child outcomes of interest are child language skills, measured using the NIH Baby Toolbox administered to all children in the study and the Receptive One Word Picture Vocabulary Test (ROWPVT) administered to children 24+ months of age in the study. We will test whether our intervention moved child language skills post-intervention for the treatment group relative to the control group.
Secondary child outcomes of interest include parent reports of child language development, non-language assessment scores, long-term child outcomes, and language measures obtained from additional processing of the audio recordings. Parent reports of child language development include the MacArthur-Bates Communicative Development Inventories (MCDI). Non-language assessments include social-emotional skills measured by the ASQ:SE, executive function skills measured by NIH Baby Toolbox, and numeracy skills measured by Woodcock Johnson-IV. All of these measures will be assessed at baseline, midline, and endline.
Long-term child outcomes include K-12 academic outcomes, which will depend on what data we can obtain from the different schools participating children will attend, but may include disciplinary records, high school graduation rates, and college enrollment rates.
We have allocated the first six weeks of the program to be the baseline period. It is possible, however, that language inputs may initially be higher than usual due to the Hawthorne effect. We also have staff that will be visiting sites more frequently the first two weeks for technology set-up and assessment administration. While child assessments are scheduled for only the first two weeks, we allow for make-up assessments which will likely bleed into the following two weeks. Due to these reasons, we will directly test whether language inputs (CTCs and AWCs) change significantly over the baseline period. If we find that the first 2-4 weeks of baseline have higher language inputs (p<0.1), then we will drop those weeks from our baseline period, and consider only the final two weeks of the baseline period. If not, we will include all six weeks of baseline to maximize power.
We include specifics about the model and analysis plan in the attached pre-analysis plan.