Secondary Outcomes (explanation)
Additional heterogeneity analysis of the primary outcome will be performed along the following dimensions:
- Socio-demographics: education level.
- Language-specific norms and attitudes. After the list experiment, participants report agreement/disagreement with up to three direct statements (depending on their list arm) capturing norms on the appropriateness of feminine-title use, the perceived harm of the feminine title to women in the profession, and the perceived ideological signaling of the feminine title. We expect all three to moderate treatment effects in the same (negative) direction; to limit multiple-hypothesis-testing concerns they will be aggregated into an index, imputing the item missing by design (for treated-arm respondents) with the sample mean of respondents giving the same responses on the other two items. At the end of the survey, respondents also complete the ATGIL scale. We will combine the direct items and the ATGIL into one language-attitudes index using the inverse-covariance-weighting approach of Anderson (2008).
- First-order beliefs: beliefs about the prevalence of women in each profession and about the prevalence of feminine-title use among women practitioners.
- Gender attitudes. From the gender-typing section (24 professions and sports), we compute the share of items the respondent classifies as suited to both men and women, as a measure of (absence of) gender-typing. This measure will be combined with the Modern Sexism score into one gender-attitudes index (inverse-covariance weighting).
- Misperception wedge (norm misperception). For respondents who answer the second-order belief questions at the national level (list arms 1, 3, and 4), an individual-level wedge is computed, following Bursztyn et al. (2020), as the respondent’s guess of the share of men (or, symmetrically, of women) who consider the feminine title more appropriate, minus the true share as measured in the study (from the list experiment and/or direct elicitation). The wedge is used as a moderator to examine whether the title-treatment effect varies with the extent and sign of a respondent’s misperception.
- Conformity-threshold wedge. From the conformity scenario we compute each respondent’s switching point — the prevalence level at which the recommendation shifts from the masculine to the feminine form. Combining this with the respondent’s first-order beliefs about the share of women using the feminine title in each profession, we compute a conformity wedge, given by the difference between the believed prevalence and the respondent’s conformity threshold, and use it as a moderator.
- Political orientation. We measure political-cultural orientation with the conservatism and liberalism subscales of the POLID instrument. Since we expect the two subscales to moderate treatment effects in opposite directions, we will reverse one and aggregate them into a single measure. In addition, the survey company provides a standard left–right self-placement measure.
Carry-over checks. Since most moderators are measured after the vignette block, we will first test whether the vignette treatment (T1–T3) has any carry-over effect on each Block-B variable. If such effects are detected for a variable, we will not use it as a moderator and will report it descriptively only. While we expect no such effects, on some dimensions (e.g., first-order beliefs about title use, the direct statements, or ATGIL) they cannot be excluded a priori, e.g., seeing only feminine-titled profiles may inflate beliefs about title use or trigger emotional reactions in some respondents.
Secondary experiments and additional exploratory analyses
Each secondary randomized component is analyzed separately, as informative about potential mechanisms; we do not analyze the full cross-randomization.
- Label-format experiment (gender-typing of professions and sports). Using the 24 gender-typing responses we test whether double-form labels (e.g., Avvocato/Avvocata) or field-name labels (e.g., Avvocatura) reduce gender-typing relative to masculine-only labels — overall, by category (STEM, non-STEM, sports, over the 4 vignette professions), and excluding professions with common gender nouns.
- List-experiment prevalence estimates. Comparing mean counts between the control arm (five baseline statements) and each treated arm (five baseline statements plus one sensitive statement) yields an estimate of the population share agreeing with each sensitive statement (“appropriateness”, “harm”, “signaling”) while preserving individual anonymity. Because the method protects anonymity by design, these estimates are aggregate overall or at sub-group level (or, at finest, subgroup-level).
- Social-desirability gap. Respondents who did not see a given sensitive statement in their own list answer the same statement directly (binary agree/disagree, identical wording, order of direct items randomized). The difference between the list-inferred prevalence and the directly stated prevalence estimates social-desirability bias for that attitude. Because list and direct responses on a given statement come from disjoint groups by construction, the gap is estimated at the aggregate level. Directly stated prevalences are computed pooling all eligible arms. As robustness checks, direct responses are tested for differences by list arm — assessing whether exposure to a different sensitive statement primes the direct response — and the gap is re-estimated using only control-arm respondents for the direct measure, controlling for the randomized position of the item within the direct-question block.
- Conformity-threshold scenario. We will estimate the effect of the female-share condition (25% vs 50%) on the conformity threshold.
References
Anderson, M. L. (2008). Multiple inference and gender differences in the effects of early intervention. Journal of the American Statistical Association, 103(484), 1481–1495.
Burlacu, S., Cappelletti, D., Marzadro, S., & Tondini, A. (2024). The cost of a vowel: How the gender-marked job title affects ratings of female lawyers. Economics Letters, 234, 111494.
Bursztyn, L., González, A. L., & Yanagizawa-Drott, D. (2020). Misperceived social norms: Women working outside the home in Saudi Arabia. American economic review, 110(10), 2997-3029.
Swim, J. K., Aikin, K. J., Hall, W. S., & Hunter, B. A. (1995). Sexism and racism: Old-fashioned and modern prejudices. Journal of Personality and Social Psychology, 68(2), 199–214.
Ulrich, M. (2021, December). Politische Ideologien (POLID). In GESIS-Leibniz-Institut Für Sozialwissenschaften (Zusammenstellung sozialwissenschaftlicher Items und Skalen). Online verfügbar unter https://doi. org/10.6102/zis313.
Schwarz, C., Dawideit, M., Hägemann, J., Ihme, J. M., Looft, J., & Nitschke, J. (2026). Inventory of attitude towards gender-inclusive language (ATGIL): Development and validation of a questionnaire as instrument to measure attitude towards gender-inclusive language for german speakers. European Journal of Psychology Open.