The Impact of Gender-Marked Professional Titles on Individual Preferences: Evidence from a Representative Survey Experiment

Last registered on August 04, 2026

Pre-Trial

Trial Information

General Information

Title
The Impact of Gender-Marked Professional Titles on Individual Preferences: Evidence from a Representative Survey Experiment
RCT ID
AEARCTR-0019185
Initial registration date
July 27, 2026

Initial registration date is when the trial was registered.

It corresponds to when the registration was submitted to the Registry to be reviewed for publication.

First published
August 04, 2026, 8:59 AM EDT

First published corresponds to when the trial was first made public on the Registry after being reviewed.

Locations

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information

Primary Investigator

Affiliation
FBK-IRVAPP

Other Primary Investigator(s)

PI Affiliation
FBK-IRVAPP
PI Affiliation
FBK-IRVAPP
PI Affiliation
Evaluation Lab @ Fondo per la Repubblica Digitale
PI Affiliation
University of Trento
PI Affiliation
FBK-IRVAPP

Additional Trial Information

Status
In development
Start date
2026-07-28
End date
2026-10-31
Secondary IDs
Prior work
This trial does not extend or rely on any prior RCTs.
Abstract
Gender-marked language may influence attitudes, preferences, and behavior, with tangible economic consequences. In Italian, professional titles carry grammatical gender (e.g., avvocato/avvocata for lawyers), yet the feminine form remains rarely used, especially in stereotypically male professions. A pilot survey experiment conducted by the research team in 2022 found that, all else equal, female lawyers presented with the feminine title were less likely to be contacted — a penalty equivalent to roughly 10 years less experience (Burlacu et al., 2024). This study extends that design to a representative sample of the Italian adult population (N = 3,200) and to four professions that differ in gender stereotypicality, in the prevalence of women in the profession, and in the prevalence of feminine-title use (law, engineering, architecture, medicine). Through a survey experiment combining between-subject randomization of title gender-marking with within-subject variation in the mixed-title arm, we investigate: (i) the causal impact of the gender marking of professional titles on stated preferences; (ii) the mechanisms underlying the observed effects, drawing on perceived descriptive and injunctive norms about title use, first- and second-order beliefs, conformity thresholds, and individual attitudes toward gender roles and gender-inclusive language; and (iii) the prevalence of sensitive beliefs about feminine professional titles — their appropriateness, their potential to penalize women in certain professions, and their perceived ideological connotation — elicited through a list experiment.
External Link(s)

Registration Citation

Citation
Burlacu, Sergiu et al. 2026. "The Impact of Gender-Marked Professional Titles on Individual Preferences: Evidence from a Representative Survey Experiment." AEA RCT Registry. August 04. https://doi.org/10.1257/rct.19185-1.0
Sponsors & Partners

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information
Experimental Details

Interventions

Intervention(s)
The study examines how the gender marking of Italian professional titles (e.g., avvocato/avvocata for lawyer) affects people’s stated preferences when choosing a professional to contact, across four professions (law, engineering, architecture, medicine). Respondents are randomly assigned to conditions that vary whether the female professional profiles they evaluate are presented with the masculine or the feminine form of the title. The survey also investigates the social norms, beliefs, and attitudes associated with the use of feminine professional titles in Italian, embedding several secondary randomized components (a label-format experiment, a list experiment, and a conformity-threshold scenario).
Intervention Start Date
2026-07-28
Intervention End Date
2026-10-31

Primary Outcomes

Primary Outcomes (end points)
Contact-likelihood rating: the stated likelihood of contacting each profile, on an integer 0–10 scale (0 = “not at all”, 10 = “for sure”). Each respondent provides 16 ratings (4 rounds × 4 profiles).
Primary Outcomes (explanation)
We expect negative effects of the feminine title, on average, in all professions except medicine, where the feminine title (Dottoressa) is prevalent and may be perceived as the social norm; medicine therefore serves as a “placebo” profession.
Extending Burlacu et al. (2024), we estimate at the profile-rating level (16 observations per respondent, standard errors clustered at the respondent level):
Rating = a0·Female + a1·T2 + a2·T3 + b1·Female×T2 + b2·Female×T3 + c1·Female×T3×FemTitle + d1·Age + d2·Experience + d3·QualitySignal1 + d4·QualitySignal2 + Scenario FE + Profile-order FE + Profession(round)-order FE + respondent characteristics + e
where Rating is the 0–10 contact-likelihood rating assigned to each profile; Female indicates a female profile; T1 (omitted category) is the arm in which all female profiles carry the masculine title, T2 the arm in which all carry the feminine title, and T3 the mixed arm; FemTitle indicates that the displayed title is the feminine form (non-zero only for female profiles in T3, and 1 for female profiles in T2 — hence the triple interaction is identified only within T3). Age, experience, and the two quality signals are the randomized profile characteristics.

a0 estimates the profile gender gap in ratings in T1; a1 and a2 estimate differences in male-profile ratings in T2 and T3 relative to T1; b1 replicates the feminine-title effect in Burlacu et al. (2024), comparing female profiles with the feminine title (T2) to female profiles with the masculine title (T1); b2 indicates whether the mere presence of feminine-titled profiles in the choice set affects female profiles carrying the masculine title (T3 masculine-titled women vs T1 women). The concern raised in the public discourse is that use of the feminine form affects all women in the profession, not only those who use it: if confirmed, b2 should be negative, with c1 negative as well only if there is an additional penalty on the specific profiles displaying the feminine title.

The functional form yields two hypothesis tests:
H0: c1 = b1. This compares the marginal feminine-title gradient in a mixed environment (T3, within-round contrast between the feminine- and masculine-titled female profile) with the gradient in a uniform environment (T2 vs T1): is the title gradient steeper when an individual evaluates profiles using both forms
H0: b2 + c1 = b1. Are feminine-titled women in the mixed arm (T3) rated differently from feminine-titled women in the uniform feminine arm (T2)

Following Burlacu et al. (2024), the main pre-specified dimension of heterogeneity is respondent gender. In addition, we will explore whether effects vary by the scenario drawn within each profession (randomized at the individual level): for each profession, one candidate scenario is relatively more stereotypically male, one more balanced, and one relatively more stereotypically female.
Finally, we will investigate whether effects vary across the three target professions. Architecture and law have similar shares of women and similarly low use of the feminine title, but differ in STEM status; architecture and engineering are both STEM and have similarly low feminine-title use, but differ markedly in the share of women (high in architecture). For the placebo profession — medicine — the feminine title is expected to have a non-negative effect, given its prevalence.
Since we do not have strong priors on heterogeneity separately for T2 and T3, when estimating heterogeneity we will pool profiles displaying the feminine title and estimate:
Rating = a0·Female + a1·T2 + a2·T3 + b2·Female×T3 + d·Female×FemTitle + [same controls and fixed effects as the main specification] + e
with each moderator interacted with the terms of interest.

Secondary Outcomes

Secondary Outcomes (end points)
Moderation of the primary treatment effect by pre-specified respondent characteristics and attitudes (heterogeneity analyses)

Secondary experiments to investigate potential mechanisms
- list-experiment prevalence estimates of three sensitive beliefs about feminine professional titles, and the corresponding social-desirability gaps relative to direct elicitation;
- gender-typing of professions and sports as an outcome of the label-format condition;
- recommendations and switching points in the conformity-threshold scenario, and their responses to the female-share and framing conditions;
- first- and second-order beliefs about women’s presence in the professions, feminine-title use, and the perceived appropriateness of the feminine title, including a norm-misperception wedge.
Secondary Outcomes (explanation)
Additional heterogeneity analysis of the primary outcome will be performed along the following dimensions:
- Socio-demographics: education level.
- Language-specific norms and attitudes. After the list experiment, participants report agreement/disagreement with up to three direct statements (depending on their list arm) capturing norms on the appropriateness of feminine-title use, the perceived harm of the feminine title to women in the profession, and the perceived ideological signaling of the feminine title. We expect all three to moderate treatment effects in the same (negative) direction; to limit multiple-hypothesis-testing concerns they will be aggregated into an index, imputing the item missing by design (for treated-arm respondents) with the sample mean of respondents giving the same responses on the other two items. At the end of the survey, respondents also complete the ATGIL scale. We will combine the direct items and the ATGIL into one language-attitudes index using the inverse-covariance-weighting approach of Anderson (2008).
- First-order beliefs: beliefs about the prevalence of women in each profession and about the prevalence of feminine-title use among women practitioners.
- Gender attitudes. From the gender-typing section (24 professions and sports), we compute the share of items the respondent classifies as suited to both men and women, as a measure of (absence of) gender-typing. This measure will be combined with the Modern Sexism score into one gender-attitudes index (inverse-covariance weighting).
- Misperception wedge (norm misperception). For respondents who answer the second-order belief questions at the national level (list arms 1, 3, and 4), an individual-level wedge is computed, following Bursztyn et al. (2020), as the respondent’s guess of the share of men (or, symmetrically, of women) who consider the feminine title more appropriate, minus the true share as measured in the study (from the list experiment and/or direct elicitation). The wedge is used as a moderator to examine whether the title-treatment effect varies with the extent and sign of a respondent’s misperception.
- Conformity-threshold wedge. From the conformity scenario we compute each respondent’s switching point — the prevalence level at which the recommendation shifts from the masculine to the feminine form. Combining this with the respondent’s first-order beliefs about the share of women using the feminine title in each profession, we compute a conformity wedge, given by the difference between the believed prevalence and the respondent’s conformity threshold, and use it as a moderator.
- Political orientation. We measure political-cultural orientation with the conservatism and liberalism subscales of the POLID instrument. Since we expect the two subscales to moderate treatment effects in opposite directions, we will reverse one and aggregate them into a single measure. In addition, the survey company provides a standard left–right self-placement measure.

Carry-over checks. Since most moderators are measured after the vignette block, we will first test whether the vignette treatment (T1–T3) has any carry-over effect on each Block-B variable. If such effects are detected for a variable, we will not use it as a moderator and will report it descriptively only. While we expect no such effects, on some dimensions (e.g., first-order beliefs about title use, the direct statements, or ATGIL) they cannot be excluded a priori, e.g., seeing only feminine-titled profiles may inflate beliefs about title use or trigger emotional reactions in some respondents.

Secondary experiments and additional exploratory analyses
Each secondary randomized component is analyzed separately, as informative about potential mechanisms; we do not analyze the full cross-randomization.
- Label-format experiment (gender-typing of professions and sports). Using the 24 gender-typing responses we test whether double-form labels (e.g., Avvocato/Avvocata) or field-name labels (e.g., Avvocatura) reduce gender-typing relative to masculine-only labels — overall, by category (STEM, non-STEM, sports, over the 4 vignette professions), and excluding professions with common gender nouns.
- List-experiment prevalence estimates. Comparing mean counts between the control arm (five baseline statements) and each treated arm (five baseline statements plus one sensitive statement) yields an estimate of the population share agreeing with each sensitive statement (“appropriateness”, “harm”, “signaling”) while preserving individual anonymity. Because the method protects anonymity by design, these estimates are aggregate overall or at sub-group level (or, at finest, subgroup-level).
- Social-desirability gap. Respondents who did not see a given sensitive statement in their own list answer the same statement directly (binary agree/disagree, identical wording, order of direct items randomized). The difference between the list-inferred prevalence and the directly stated prevalence estimates social-desirability bias for that attitude. Because list and direct responses on a given statement come from disjoint groups by construction, the gap is estimated at the aggregate level. Directly stated prevalences are computed pooling all eligible arms. As robustness checks, direct responses are tested for differences by list arm — assessing whether exposure to a different sensitive statement primes the direct response — and the gap is re-estimated using only control-arm respondents for the direct measure, controlling for the randomized position of the item within the direct-question block.
- Conformity-threshold scenario. We will estimate the effect of the female-share condition (25% vs 50%) on the conformity threshold.

References
Anderson, M. L. (2008). Multiple inference and gender differences in the effects of early intervention. Journal of the American Statistical Association, 103(484), 1481–1495.
Burlacu, S., Cappelletti, D., Marzadro, S., & Tondini, A. (2024). The cost of a vowel: How the gender-marked job title affects ratings of female lawyers. Economics Letters, 234, 111494.
Bursztyn, L., González, A. L., & Yanagizawa-Drott, D. (2020). Misperceived social norms: Women working outside the home in Saudi Arabia. American economic review, 110(10), 2997-3029.
Swim, J. K., Aikin, K. J., Hall, W. S., & Hunter, B. A. (1995). Sexism and racism: Old-fashioned and modern prejudices. Journal of Personality and Social Psychology, 68(2), 199–214.
Ulrich, M. (2021, December). Politische Ideologien (POLID). In GESIS-Leibniz-Institut Für Sozialwissenschaften (Zusammenstellung sozialwissenschaftlicher Items und Skalen). Online verfügbar unter https://doi. org/10.6102/zis313.
Schwarz, C., Dawideit, M., Hägemann, J., Ihme, J. M., Looft, J., & Nitschke, J. (2026). Inventory of attitude towards gender-inclusive language (ATGIL): Development and validation of a questionnaire as instrument to measure attitude towards gender-inclusive language for german speakers. European Journal of Psychology Open.


Experimental Design

Experimental Design
The study is an online survey experiment on a representative sample of Italian adults (SWG CAWI panel; completion time approximately 15 minutes), comprising two blocks. In Block A, respondents complete 4 rounds of profile evaluation, one per profession (law, engineering, architecture, medicine); for each profession, one service-need scenario is drawn at random from three candidates. In each round they view 4 hypothetical professional profiles (2 male, 2 female) and rate each profile on the likelihood of contacting them on a 0–10 scale. Profiles display the professional title with first and last name, year of birth, year of professional-register enrollment, and two quality signals. The gender marking of professional titles is randomized between subjects: throughout the 4 rounds, some respondents see female profiles titled exclusively with the masculine form, others exclusively with the feminine form, and others encounter both forms (one female profile with each form in every round); male profiles are always presented with the masculine form.
In Block B, respondents complete attitudinal and belief-elicitation sections: gender-typing of professions and sports (with randomized label formats), a list experiment on attitudes toward feminine professional titles, first- and second-order belief elicitation, direct attitude questions, a conformity-threshold scenario with randomized conditions, and validated attitudinal scales (POLID, Ulrich et al., 2021; Modern Sexism, Swim et al., 1995; ATGIL, Schwarz et al., 2026). Section order is fixed and designed to limit priming; item and response-option orders are randomized. All randomizations are assigned at the individual level, independently of one another.

Experimental Design Details
Not available
Randomization Method
Randomization at the individual respondent level, implemented in software (oTree) at session start. All treatments and randomized orderings are drawn independently.
Randomization Unit
Respondent level
Was the treatment clustered?
No

Experiment Characteristics

Sample size: planned number of clusters
3,200 individuals (randomization is at the individual level; no clustering): adults (18+) residing in Italy, members of the SWG CAWI panel, representative by gender, age group, and Istat geographic macro-area.
Sample size: planned number of observations
3,200 respondents. For the vignette analysis, each respondent rates 16 profiles, yielding 51,200 profile ratings
Sample size (or number of clusters) by treatment arms
Title treatment (vignette): ~960 respondents in T1 (all female profiles with the masculine title), ~960 in T2 (all with the feminine title), ~1,280 in T3 (mixed titles).
List experiment (cross-randomized): ~960 control, ~960 “appropriateness”, ~640 “harm”, ~640 “signaling”).
Label condition (gender-typing section, cross-randomized): ~1,067 respondents per arm (masculine-only, double form, field name).
Conformity-threshold scenario (cross-randomized 2×2): ~800 respondents per cell (female share 25% vs 50% × feminine- vs masculine-anchored framing); ~1,600 respondents per level on each margin.
Minimum detectable effect size for main outcomes (accounting for sample design and clustering)
Assuming a mean rating of 6.5, a SD of 2.5 and an ICC of 0.3, and considering only the 12 ratings for the 3 professions of interest, the MDE relative to a baseline group (T1) is roughly 0.19 points, or 0.076 SD
IRB

Institutional Review Boards (IRBs)

IRB Name
IRB Approval Date
IRB Approval Number