Experimental Design Details
Item list
The experiment includes 25 items across three mutually exclusive categories: 10 STEM occupations, 10 non-STEM occupations, and 5 sports. Within each category, items are classified as grammatically declinable (19 items) or grammatically invariable (6 items) in Italian.
- STEM occupations (10):
- Declinable in Italian (8): Computer programmer, Engineer, Architect, Biologist, Astronomer, Researcher, Surgeon, Nurse
- Invariable in Italian (2): Data analyst, Pharmacist
- Non-STEM occupations (10):
- Declinable in Italian (7): Psychologist, Police officer, Lawyer, Photographer, Pastry chef, Translator, Farmer
- Invariable in Italian (3): School principal, Social worker, Journalist
- Sports (5):
- Declinable in Italian (4): Skier, Football player, Dancer, climber
- Invariable in Italian (1): Cyclist
Hypotheses:
H1: Participants in the pair form condition will show more egalitarian perceptions than participants in the generic masculine condition (positive coefficient on D_pair).
H2: Participants in the professional domain condition will show intermediate perceptions between the generic masculine and pair form conditions (positive coefficient on D_domain, smaller than D_pair).
H3: The effect of language condition will be stronger for STEM occupations and sports than for non-STEM occupations, tested as positive interaction terms between language condition dummies and D_STEM and D_Sport.
H4: In the pair form condition, inclusive priming will extend to grammatically invariable items, increasing the probability of a “both” response relative to the generic masculine condition. Tested in a separate regression on invariable items only.
H5: The effect of language condition will be non-linearly moderated by pre-existing gender stereotype levels, measured at baseline prior to both SPARKLE treatment assignment and language randomization. Specifically, we hypothesize an inverted U-shaped relationship: the effect of inclusive language conditions (pair form and professional domain) relative to the generic masculine will be strongest for students with intermediate stereotype levels, and attenuated for students at both ends of the stereotype distribution. Students with very low stereotype levels (highly egalitarian) are expected to show a ceiling effect; students with very strong (negative) stereotype levels are expected to show resistance to inclusive linguistic cues. This is tested by including both a linear and a quadratic interaction term between language condition dummies and the baseline stereotype scale. The non-linear moderation hypothesis predicts a significant quadratic interaction term, with the effect of language condition being strongest for students with intermediate stereotype levels. Given the much lower levels of stereotypes among female students at baseline, the U shape may not be observed.
Estimation strategy:
Primary analyses (H1, H2, and H5) are conducted on the control group of the SPARKLE RCT only, using the sample of grammatically declinable items (19 items). The outcome is a binary dummy equal to 1 if the participant responded "both" for a given item. The model is an OLS regression with standard errors clustered at the individual level, estimated separately for boys and girls. All models include language condition dummies (D_pair and D_domain, with generic masculine as reference), item category dummies (D_STEM and D_Sport, with non-STEM as reference), and individual-level controls (e.g., cognitive ability, math grade, year of birth, parental education, language spoken at home). H5 adds the baseline stereotype scale, its square, and their interactions with language condition dummies. A separate primary analysis (H4) uses the same model structure estimated on the 6 grammatically invariable items. The exploratory secondary analysis uses the full sample (treated and control), adding SPARKLE treatment assignment and the interaction terms D_pair × SPARKLE and D_domain × SPARKLE.