Back to History

Fields Changed

Registration

Field Before After
Abstract A growing body of evidence demonstrates that laws and institutions shape social norms — not only through enforcement but through their signalling function: communicating what is socially acceptable and legitimising or delegitimising group identities in the public sphere. This signalling effect is theorised to be especially powerful when institutional signals are reinforced by religious authority, yet causal evidence from low-institutional-legitimacy, high-religious-salience contexts remains scarce. In contexts such as Pakistan, where the credibility of secular state institutions is constrained and Islamic authority operates as the primary mechanism of norm legitimation, understanding how religious framing modulates the effect of institutional signals on public behaviour is both theoretically important and practically urgent. This study provides the first large-scale randomised experimental test of this mechanism, exploiting the live and unresolved legal context of Pakistan's Transgender Persons (Protection of Rights) Act 2018 — currently under Supreme Court review following its partial annulment by the Federal Shariat Court on Islamic grounds in 2023. Approximately 2,000–3,000 adults across urban localities in Lahore are randomly assigned to one of five conditions: a control condition receiving no institutional signal, or one of four vignette treatments describing an anticipated 2027 Supreme Court ruling either upholding or annulling the Act, with or without explicit religious framing drawn from real institutional discourse — the Council of Islamic Ideology's prior endorsement of the Act as consistent with Islamic principles, or the Federal Shariat Court's Quranic jurisprudential reasoning against it. The study measures whether and how these signals shift: perceived descriptive and injunctive social norms regarding transgender inclusion; personal attitudes toward transgender individuals across social, economic, and educational domains; pluralistic ignorance — the gap between personal attitudes and perceived social norms; labour market perceptions including perceived employability of transgender individuals; and real behavioural outcomes including incentivised donation allocation, voluntary enrolment in transgender-staffed household and tutoring services, and willingness to publicly associate with transgender inclusion through information sharing. A difference-in-differences/ANCOVA design on pre- and post-treatment institutional trust measures tests whether judicial signals shift the very legitimacy of the institutions sending them. Heterogeneity analysis examines whether religious framing amplifies institutional signals differentially for high-religiosity respondents, whether this amplification is symmetric across inclusive and exclusionary signals, and whether prior knowledge of the Act moderates treatment effects through belief updating. The study contributes to three literatures: the economics of social norm formation and institutional signalling; the political economy of religion and governance in Muslim-majority societies; and the economics of discrimination and labour market exclusion of stigmatised minorities. Findings have direct policy relevance for the design of judicial communication strategies, civil society advocacy, and inclusion programming in contexts where conventional rights-based institutional legitimacy is weak and religious authority functions as the primary vehicle for norm change. A growing body of evidence demonstrates that laws and institutions shape social norms — not only through enforcement but through their signalling function: communicating what is socially acceptable and legitimising or delegitimising group identities in the public sphere. This signalling effect is theorised to be especially powerful when institutional signals are reinforced by religious authority, yet causal evidence from low-institutional-legitimacy, high-religious-salience contexts remains scarce. In contexts such as Pakistan, where the credibility of secular state institutions is constrained and Islamic authority operates as the primary mechanism of norm legitimation, understanding how religious framing modulates the effect of institutional signals on public behaviour is both theoretically important and practically urgent. This study provides the first large-scale randomised experimental test of this mechanism, exploiting the live and unresolved legal context of Pakistan's Transgender Persons (Protection of Rights) Act 2018 — currently under Supreme Court review following its partial annulment by the Federal Shariat Court on Islamic grounds in 2023. Approximately 3,500 adults, drawn from a density-stratified, population-representative sample of urban Lahore, are randomly assigned to one of five conditions: a control condition receiving no institutional signal, or one of four vignette treatments describing an anticipated Supreme Court ruling either upholding or annulling the Act, with or without explicit religious framing drawn from real institutional discourse — the Council of Islamic Ideology's prior endorsement of the Act as consistent with Islamic principles, or the Federal Shariat Court's Quranic jurisprudential reasoning against it. The study measures whether and how these signals shift: perceived descriptive and injunctive social norms regarding transgender inclusion; personal attitudes toward transgender individuals across social, economic, and educational domains; pluralistic ignorance — the gap between personal attitudes and perceived social norms; labour market perceptions including perceived employability of transgender individuals; and real behavioural outcomes including incentivised donation allocation, voluntary enrolment in transgender-staffed household and tutoring services, and willingness to publicly associate with transgender inclusion through information sharing. A difference-in-differences/ANCOVA design on pre- and post-treatment institutional trust measures tests whether judicial signals shift the very legitimacy of the institutions sending them, alongside a parallel measure of the gap between personal and perceived-community trust in the Supreme Court, included as a validity check on social desirability bias given the study's implementation through a government-affiliated fieldwork partner. Heterogeneity analysis examines whether religious framing amplifies institutional signals differentially for high-religiosity respondents, whether this amplification is symmetric across inclusive and exclusionary signals, and whether prior knowledge of the Act moderates treatment effects through belief updating. The study contributes to three literatures: the economics of social norm formation and institutional signalling; the political economy of religion and governance in Muslim-majority societies; and the economics of discrimination and labour market exclusion of stigmatised minorities. Findings have direct policy relevance for the design of judicial communication strategies, civil society advocacy, and inclusion programming in contexts where conventional rights-based institutional legitimacy is weak and religious authority functions as the primary vehicle for norm change.
Last Published August 19, 2026 09:13 PM August 24, 2026 09:55 AM
Intervention (Public) Participants are randomly assigned to one of five experimental conditions with equal probability (0.20 each). All participants first complete a demographics and baseline moderators block measuring institutional trust, religiosity, social norm sensitivity, and prior contact with the transgender community. The control condition (C) receives no information about the Supreme Court case. Four treatment conditions present participants with a vignette describing an anticipated 2027 Supreme Court ruling on the Transgender Persons (Protection of Rights) Act 2018: T1 — Positive institutional signal without religious framing: the Supreme Court is expected to uphold the Act and maintain all existing protections for transgender individuals. T2 — Negative institutional signal without religious framing: the Supreme Court is expected to strike down the Act and revoke existing protections. T3 — Positive institutional signal with Islamic religious framing: same as T1, with additional framing citing Islamic principles of justice, compassion, and human dignity, consistent with Islamic teachings in the Holy Quran and Sunnah. T4 — Negative institutional signal with Islamic religious framing: same as T2, with additional framing citing the Federal Shariat Court's 2023 reasoning that the Act contradicts Quranic injunctions on the divine fixity of gender, drawing on Surah An-Nisa 4:11-12. All vignettes are read aloud by trained enumerators and are designed to reflect real institutional discourse previously used by Pakistani state and religious bodies. A manipulation check immediately following the vignette measures perceived likelihood of the ruling to test vignette credibility. Outcome measures follow in the order: social norm perception, personal attitudes, behavioural measures, post-treatment institutional trust. Participants are randomly assigned to one of five experimental conditions with equal probability (0.20 each). All participants first complete a demographics and baseline moderators block measuring institutional trust, religiosity, social norm sensitivity, and prior contact with the transgender community. The control condition (C) receives no information about the Supreme Court case. Four treatment conditions present participants with a vignette describing an anticipated Supreme Court ruling on the Transgender Persons (Protection of Rights) Act 2018: T1 — Positive institutional signal without religious framing: the Supreme Court is expected to uphold the Act and maintain all existing protections for transgender individuals. T2 — Negative institutional signal without religious framing: the Supreme Court is expected to strike down the Act and revoke existing protections. T3 — Positive institutional signal with Islamic religious framing: same as T1, with additional framing citing Islamic principles of justice, compassion, and human dignity, consistent with Islamic teachings in the Holy Quran and Sunnah. T4 — Negative institutional signal with Islamic religious framing: same as T2, with additional framing citing the Federal Shariat Court's 2023 reasoning that the Act contradicts Quranic injunctions on the divine fixity of gender, drawing on Surah An-Nisa 4:11-12. All vignettes are read aloud by trained enumerators and are designed to reflect real institutional discourse previously used by Pakistani state and religious bodies. A manipulation check immediately following the vignette measures the perceived likelihood of the ruling to test vignette credibility. Outcome measures follow in the order: social norm perception, personal attitudes, behavioural measures, post-treatment institutional trust.
Primary Outcomes (End Points) Perceived descriptive social norms index: mean response across eleven items measuring what most people in the respondent's community are believed to feel regarding social proximity, labour market inclusion, and educational inclusion of transgender individuals (5-point agree/disagree scale) Personal attitudes index: mean response across fourteen items measuring respondent's own comfort with and support for social proximity, labour market inclusion, and educational inclusion of transgender individuals (5-point agree/disagree scale) Donation allocation: amount donated (PKR 0–300) from guaranteed participation payment to a transgender employment NGO (Gender Rights Watch) Service enrolment — cleaning: binary indicator of whether the respondent signs up for free household cleaning services delivered by trained transgender employees Service enrolment — tutoring: binary indicator of whether respondents with school-age children enrol in free tutoring services delivered by qualified transgender teachers (conditional on having school-age children in household) Perceived descriptive social norms index: mean response across eleven items measuring what most people in the respondent's area, or in Pakistan, are believed to feel or believe regarding social proximity, labour market inclusion, and educational inclusion of transgender individuals. Ten items use a 0–10 count-based response format ("out of 10 people in your area, how many...") and one item (measuring perceived national-level trend) uses a 5-point Likert scale, reflecting the difference in reference group scope. Personal attitudes index: mean response across sixteen items measuring respondent's own comfort with and support for social proximity, labour market inclusion, and educational inclusion of transgender individuals (5-point agree/disagree scale). Donation allocation: amount donated (PKR 0–150, in discrete increments of 0/10/20/50/100/150) from guaranteed participation payment to a transgender employment program. Service enrolment — cleaning: binary indicator of whether the respondent signs up for free household cleaning and household help services delivered by trained transgender employees. Service enrolment — tutoring: binary indicator of whether respondents with school-age children enrol in free tutoring services delivered by qualified transgender teachers (conditional on having school-age children in household).
Primary Outcomes (Explanation) The perceived norms index and personal attitudes index will each be constructed as Anderson (2008) summary indices — standardised z-scores averaged across constituent items — to address multiple outcomes within each family and to maximise power. Items within each index will be standardised to mean zero and unit standard deviation using the control group mean and standard deviation before being averaged. Reverse-coded items (two per block) will be recoded before standardisation so that higher values consistently indicate greater acceptance of transgender inclusion. Cronbach's alpha will be reported for each index. The donation measure is a continuous variable (PKR 0 to 250) and will be analysed both in levels and as a binary indicator (donated any amount versus none). Enrolment measures are binary (yes/no) and will be analysed using linear probability models and logistic regression as robustness checks. The perceived norms index and personal attitudes index will each be constructed as Anderson (2008) summary indices — standardised z-scores averaged across constituent items — to address multiple outcomes within each family and to maximise power. Because index items are collected on different raw response scales (0–10 count format for area-level norm items, 5-point Likert for national-level norm items and all personal attitude items), each item is standardised against its own distribution (mean and standard deviation, using the control group) before being averaged into its respective index; this standardisation is what permits valid combination of items collected on different raw scales. Reverse-coded items will be recoded before standardisation so that higher values consistently indicate greater acceptance of transgender inclusion. Cronbach's alpha will be reported for each index. The donation measure is a discrete variable (PKR 0 to 150) and will be analysed both in levels and as a binary indicator (donated any amount versus none). Enrolment measures are binary (yes/no) and will be analysed using linear probability models and logistic regression as robustness checks.
Experimental Design (Public) The study uses a between-subjects randomised vignette experiment embedded in a household survey. The design has five arms — one control and four treatments — with equal probability of assignment (0.20 per arm). Assignment is individual-level, implemented through Qualtrics/Kobo randomiser software administered by trained enumerators on tablets during face-to-face household visits. The four treatment conditions follow a 2x2 factorial structure crossing two dimensions: policy direction (positive ruling upholding the Transgender Persons Act versus negative ruling striking it down) and religious framing (absent versus present). The control condition receives no information about the Supreme Court case. The survey flow is: demographics and baseline moderators (including pre-treatment institutional trust and religiosity measures) → prior knowledge questions → randomisation → vignette or control → manipulation check → outcome measures (social norm perception, personal attitudes, behavioural measures) → post-treatment institutional trust measures. Household sampling uses geographically clustered systematic sampling within selected Union Councils in urban Lahore. Streets within each UC are mapped and randomised using QGIS. Within selected streets, every third household is approached starting from a randomly dropped GPS pin. Within each household, one adult is randomly selected using a pre-generated random number assignment protocol. Enumerators follow a callback protocol for unavailable respondents. The study uses a between-subjects randomised vignette experiment embedded in a household survey. The design has five arms — one control and four treatments — with equal probability of assignment (0.20 per arm). Assignment is individual-level, implemented through the KoboToolbox survey instrument administered by trained enumerators on tablets during face-to-face household visits. The four treatment conditions follow a 2x2 factorial structure crossing two dimensions: policy direction (positive ruling upholding the Transgender Persons Act versus negative ruling striking it down) and religious framing (absent versus present). The control condition receives no information about the Supreme Court case. The survey flow is: demographics and baseline moderators (including pre-treatment institutional trust, perceived institutional trust, and religiosity measures) → prior knowledge questions → randomisation → vignette or control → manipulation check → outcome measures (social norm perception, personal attitudes, behavioural measures) → post-treatment institutional trust measures. Household sampling uses a density-stratified sampling frame across 68 Union Councils in urban Lahore. UC-level population density (WorldPop 2025, 100-metre resolution) was used to classify each UC into one of two sampling procedures based on an empirically-derived threshold of 10,900 persons per km². For UCs at or above this threshold (47 of 68), reflecting sufficiently dense residential development, starting points were generated via simple random sampling within the UC boundary polygon. For UCs below this threshold (21 of 68), where administrative "Urban" classification was found to include substantial non-residential land, starting points were generated via constrained random sampling within a 150-metre radius of satellite-imagery-verified residential reference points (three per UC), sampled without replacement to preserve spatial diversity. Five starting points were generated per UC (ten for two double-weighted UCs), for 350 starting points in total. From each starting point, enumerators follow a systematic right-hand-rule walk protocol, approaching every third household, with a defined substitution protocol for non-response and a single callback permitted before substitution. Within each household, one adult is randomly selected using a random draw generated within the survey software. Independently of treatment assignment, respondents are randomly assigned to one of two response elicitation modes for all scale-based items (10% verbal, 90% non-verbal card-based response), included as a measurement validity check and not crossed with treatment arm in the analysis of primary outcomes.
Randomization Method Individual-level randomisation implemented by Qualtrics/Kobo survey software at the point of survey administration. Each tablet-based survey session is independently randomised with equal probability (0.20) to one of five arms using Qualtrics'/Kobo's built-in randomiser function with 'evenly present elements' option enabled. For offline administration, a pre-generated randomisation list created in Stata using a documented random seed is used, with treatment assignment entered by enumerators at the start of each survey. The randomisation seed and full assignment list will be archived and made available in the replication package. Treatment assignment is generated within the KoboToolbox survey instrument at the start of each interview using the formula once(int(random()*5)+1), which draws a value between 1 and 5 with equal probability (0.20 each) and fixes it for the remainder of that respondent's session. Assignment is independently generated per submission. Response mode (verbal vs. card-based) is assigned independently using an analogous formula with unequal probability weighting (0.10 / 0.90). The full de-identified assignment log for both randomisations will be archived and made available in the replication package.
Planned Number of Clusters 50-60 Union Councils in urban Lahore (the unit used for standard error clustering in estimation, not the randomisation unit) 68 unique Union Councils in urban Lahore (70 sampling slots; two UCs are double-weighted and receive double the standard number of starting points). This is the unit used for standard error clustering in estimation, not the randomisation unit.
Planned Number of Observations 2,000–3,000 individual adult respondents 3,500 individual adult respondents is the target completed and QA-valid analysed sample. Household-level non-response (refusal to participate at the point of contact) is addressed in real time through the field substitution protocol described under Experimental Design, and does not directly enter this calculation. Separately, based on an anticipated post-hoc quality-assurance invalidation rate of approximately 20% — covering interviews terminated before completion, failed telephone back-checks, or surveys flagged during audio review — enumerators will aim to complete approximately 13 field interviews per starting point (rather than the nominal 10) across 350 starting points, in order to net 10 QA-valid interviews per point and reach the target of 3,500 valid observations overall. If the realised invalidation rate is lower than anticipated, the final valid sample may exceed 3,500.
Sample size (or number of clusters) by treatment arms Approximately 400–600 respondents per arm across five arms (Control: 400–600; T1 Positive no religion: 400–600; T2 Negative no religion: 400–600; T3 Positive with religion: 400–600; T4 Negative with religion: 400–600). Equal allocation across arms. 700 respondents per arm across five arms (Control: 700; T1 Positive, no religious framing: 700; T2 Negative, no religious framing: 700; T3 Positive, with religious framing: 700; T4 Negative, with religious framing: 700). Equal allocation across arms, total N = 3,500.
Power calculation: Minimum Detectable Effect Size for Main Outcomes Based on effect sizes reported in Tankard and Paluck (2017) for comparable institutional signalling vignette experiments — standardised effect sizes of d=0.15 to d=0.25 on norm perception and attitude outcomes — the following power calculations apply: At n=400 per arm, two-arm comparison (any treatment vs control), two-tailed alpha=0.05, power=0.80: minimum detectable effect size d=0.20. At n=600 per arm, same parameters: minimum detectable effect size d=0.16. The study is adequately powered to detect main treatment effects at the lower bound of effect sizes observed in comparable experiments. For pre-registered heterogeneity analyses (religiosity and prior knowledge moderators), splitting each arm on a binary moderator produces subgroups of approximately 200–300 respondents per subgroup. Minimum detectable interaction effect sizes at 80% power for n=300 per subgroup: d=0.23. The study is therefore powered for heterogeneity tests of moderate to large interaction effects but not for small interactions, which will be reported as directional with appropriate confidence intervals. For behavioural outcomes (donation allocation, service enrolment), power calculations assume a base rate of 30% enrolment in the control group based on comparable NGO opt-in experiments in South Asia. At n=400 per arm, the study detects differences in proportions of approximately 7–8 percentage points at 80% power. At n=600 per arm, approximately 6 percentage points. Based on effect sizes reported in Tankard and Paluck (2017) for comparable institutional signalling vignette experiments — standardised effect sizes of d=0.15 to d=0.25 on norm perception and attitude outcomes — the following power calculations apply at the study's target sample size: At n=700 per arm, two-arm comparison (any treatment vs. control), two-tailed alpha=0.05, power=0.80: minimum detectable effect size d=0.15. The study is adequately powered to detect main treatment effects at or below the lower bound of effect sizes observed in comparable experiments. For pre-registered heterogeneity analyses (religiosity and prior knowledge moderators), splitting each arm on a binary moderator produces subgroups of approximately 350 respondents. Minimum detectable interaction effect sizes at 80% power for n=350 per subgroup: d=0.21. The study is powered for heterogeneity tests of moderate interaction effects; smaller interactions will be reported as directional with appropriate confidence intervals. For behavioural outcomes (donation allocation, service enrolment), power calculations assume a base rate of 30% enrolment in the control group based on comparable NGO opt-in experiments in South Asia. At n=700 per arm, the study detects differences in proportions of approximately 6.9 percentage points at 80% power. Because the field protocol targets at least 3,500 QA-valid observations and may yield a modestly larger realised sample depending on the actual invalidation rate, all power calculations above should be read as conservative upper bounds on minimum detectable effect size; a larger realised sample would tighten the main-effect MDE somewhat more than the heterogeneity MDE, reflecting the structurally lower sensitivity of interaction-effect power to sample size increases.
Secondary Outcomes (End Points) Injunctive social norms index: mean response across two items measuring what most people believe should happen regarding transgender rights and legal protections Pluralistic ignorance measure: within-respondent gap between personal attitude score and perceived descriptive norm score — a positive gap indicates the respondent is more accepting than they believe their community to be Labour market perceptions: two items measuring perceived employability of transgender individuals in professional versus informal labour Post-treatment institutional trust index: Anderson summary index of five items measuring trust in the Supreme Court and government, repeated post-treatment to enable difference-in-differences estimation of whether judicial signals shift institutional legitimacy Religious identity activation: single post-treatment item measuring the extent to which religious beliefs influenced the respondent's views during the survey — mechanism test for religious framing effects Leaflet acceptance: binary indicator of whether respondent accepts a leaflet about transgender employment opportunities for sharing with others — measure of willingness to be publicly associated with transgender inclusion Manipulation check: perceived likelihood of the ruling (1–5 scale) — used as a moderator to test whether treatment effects are larger among respondents who found the vignette credible Injunctive social norms index: mean response across three items measuring what most people believe should happen, or what transgender individuals are believed to deserve, regarding rights and legal protections. Two items use a 0–10 count-based response format (area-level reference group) and one item uses a 5-point Likert scale (national-level reference group). Pluralistic ignorance measure: within-respondent gap between personal attitude score and perceived descriptive norm score — a positive gap indicates the respondent is more accepting than they believe their community to be. Labour market perceptions: two items measuring perceived employability of transgender individuals in professional versus informal labour. Post-treatment institutional trust index: Anderson summary index of five items measuring trust in the Supreme Court and government, repeated post-treatment to enable estimation of whether judicial signals shift institutional legitimacy. Institutional trust misperception (preference-falsification gap): within-respondent gap between personal trust in the Supreme Court and perceived community trust in the Supreme Court, measured pre- and post-treatment, and benchmarked separately at Union Council, tehsil, and city level — included as a validity check on social desirability bias in institutional trust reporting, given the general political sensitivity in Pakistan around expressing distrust of the judiciary to an unfamiliar interviewer, rather than as a substantive pluralistic ignorance finding. Religious identity activation: single post-treatment item measuring the extent to which religious beliefs influenced the respondent's views during the survey — mechanism test for religious framing effects. Leaflet acceptance: binary indicator of whether respondent accepts a leaflet about transgender employment opportunities for sharing with others — measure of willingness to be publicly associated with transgender inclusion. Response-mode validity check: comparison of attitude and norm-perception index scores between verbal and card-based (non-verbal) response conditions, pooled across treatment arms, to assess whether response format affects reporting on sensitive items. Service enrolment persistence (exploratory, subject to feasibility): a brief follow-up contact with respondents who enrolled in either the cleaning/household help or tutoring service, conducted several weeks after the baseline survey, to assess whether the initial enrolment decision translated into actual service uptake, and whether this persistence differs across treatment arms. This follow-up is conditional on field operational capacity and is not itself a primary registered outcome; if conducted, results will be reported as exploratory and will not be used to revise conclusions drawn from the primary pre-registered enrolment measures.
Secondary Outcomes (Explanation) Pluralistic ignorance will be computed as the individual-level difference between the respondent's personal attitudes index score and their perceived descriptive norms index score. A positive value indicates the respondent is more accepting than they perceive their community to be, consistent with pluralistic ignorance in the inclusion direction. Treatment effects on pluralistic ignorance will be estimated both directly (as a constructed outcome) and by comparing treatment effects on the two underlying indices separately, allowing identification of whether signals primarily shift perceived norms, personal attitudes, or both. The institutional trust DiD will be estimated as: (post-treatment trust score minus pre-treatment trust score) compared across treatment arms, with the control group change providing the counterfactual for survey-induced drift in trust scores. Pluralistic ignorance will be computed as the individual-level difference between the respondent's personal attitudes index score and their perceived descriptive norms index score. A positive value indicates the respondent is more accepting than they perceive their community to be, consistent with pluralistic ignorance in the inclusion direction. Treatment effects on pluralistic ignorance will be estimated both directly (as a constructed outcome) and by comparing treatment effects on the two underlying indices separately, allowing identification of whether signals primarily shift perceived norms, personal attitudes, or both. The institutional trust misperception gap is constructed identically to the pluralistic ignorance measure above (personal score minus perceived-community score, both standardised before differencing) but is interpreted in the opposite direction: a positive gap here is expected given the general sensitivity around openly expressing distrust of state institutions in the current political climate, and serves as a check on desirability bias in institutional trust reporting rather than as evidence about genuine community norms. Perceived-community trust and perceptions ("میرا علاقہ" / "my area" in the Urdu instrument) are benchmarked against the average personal trust score among Control-arm respondents at three levels of geographic aggregation — Union Council, tehsil, and city — to assess whether the misperception gap is sensitive to the scale at which "community" is defined, given known limits on statistical precision at the Union Council level with the achieved per-arm sample size. The post-treatment institutional trust index will be analysed using both a difference-in-differences specification and an ANCOVA specification (post-treatment score regressed on treatment arm indicators, controlling for the pre-treatment score as a covariate). Because treatment is individually randomised, both approaches yield unbiased estimates of the treatment effect; however, ANCOVA is expected to provide a more statistically powerful and more precise test of whether judicial signals shift institutional legitimacy, consistent with established methodological findings that ANCOVA dominates change-score analysis in randomised designs (Van Breukelen, 2006). The ANCOVA specification will serve as the primary specification for this outcome, with the DiD estimate reported as a robustness check. Independently of treatment arm, 10% of respondents complete the survey using standard verbal response elicitation while 90% use a non-verbal, card-based method in which the respondent points to a response rather than stating it aloud. This assignment is orthogonal to treatment and is not crossed with treatment arm in any planned analysis of primary outcomes; the comparison is included solely to assess measurement validity. The service enrolment persistence check, if conducted, addresses the distinction between an immediate enrolment decision — a "flow" outcome potentially sensitive to short-term emotional or social-desirability responses to the vignette — and durable behavioural follow-through, better evidence of a genuine change in willingness to engage with transgender-provided services. Because this follow-up depends on field team capacity and resources beyond the primary data collection window, it is registered here as a conditional, exploratory extension rather than a primary outcome.
Back to top