Experimental Design Details
General structure of the experiment
The experiment will run online on Prolific. After obtaining informed consent, participants will first report details about their most recent rental experience. The main task consists of 33 rounds divided into three budget blocks. In each round, participants are presented with two fictitious properties and must choose the one they would prefer to rent. To encourage truthful revelation of preferences and mitigate social desirability bias, participants are incentivised with bonus payments based on how closely their choices align with the modal choice (the option most frequently selected by other participants). The experiment concludes with a post-experimental survey comprising an Implicit Association Test to measure implicit biases and standard demographic questions.
Treatments
Our experimental design consists of three between-participant treatments. In all treatments, we use a within-subjects orthogonal design where participants evaluate 28 pairs of properties across three primary dimensions.
•Treatment 1 (Quantity): We vary the host race (minority/non-minority), host gender (man/woman), and review quantity (low/high, keeping the quality of reviews fixed).
•Treatment 2 (Positive Informativeness): We vary the host race, host gender, and the informativeness of reviews (low/high, keeping the number of reviews fixed) when all reviews are positive.
•Treatment 3 (Negative Informativeness): We vary the host race, host gender, and the informativeness of reviews (low/high, keeping the number of reviews fixed) when one of the reviews is negative.
Participants are randomly assigned to one of the three treatments. To prevent participants from seeing the exact same property twice while ensuring that all combinations of traits are tested against each other, the 28 experimental rounds are constructed using a predefined "property map". To achieve perfect counterbalancing, participants within each treatment are randomly assigned to one of 56 distinct participant types. This rotation ensures that property attributes are orthogonal to the round number and budget tier, minimising order effects.
Hypotheses and Analysis of Main Effects
We will estimate a Linear Probability Model (LPM) using the stacked (long) data format. We estimate the following specification separately for each treatment arm:
Y_ijt = α + β_1 Black_j + β_2 Man_j + β_3 LowRev_j + β_12 (Black_j * Man_j) + β_13 (Black_j * LowRev_j) + β_23 (Man_j * LowRev_j) + β_123 (Black_j * Man_j * LowRev_j) + δ_t + ε_ijt
Where:
• Y_ijt is a binary indicator equal to 1 if participant i chose property j in round t, and 0 otherwise.
• Black_j, Man_j, and LowRev_j are dummy variables for the property attributes. The baseline category is the "privileged" profile (White, Woman, High Reviews), meaning coefficients represent penalties relative to this ideal.
• δ_t are block fixed effects (Low, Mid, High budgets).
• Standard errors (ε_ijt) are clustered at the participant level i.
This fully saturated specification allows us to test the following hypotheses:
1. H1–H3 (Main Effects): We test for the significance of β_1 (Race), β_2 (Gender), and β_3 (Reviews) to determine the baseline penalties or premiums associated with each attribute.
2. H4–H6 (Two-Way Interactions): We test the coefficients β12 (Race * Gender), β_13 (Race * Reviews), and β_23 (Gender * Reviews) to identify if penalties are additive or compounding (e.g., if the racial penalty is significantly different for men vs. women).
3. H7 (Three-Way Interaction): We test β_123 (Race * Gender * Reviews) to determine if the effect of reviews on racial discrimination depends simultaneously on the host's gender.
Exploratory analysis
Heterogeneity (In-group Bias / Homophily):
We will investigate whether the estimated effects vary by participant characteristics. Because the outcome is a forced choice with a fixed mean of 0.5, participant characteristics cannot be included as simple controls. Instead, we include them as moderators by interacting participant demographics (Z_i) with property attributes (X_j):
Y_ijt = α + β X_j + λ (X_j * Z_i) + ε_ijt
Specific hypotheses include:
• Gender Homophily: We test if λ_{WomanHost * WomanRater} > 0, indicating that female participants are relatively more likely to choose female hosts.
• Racial In-Group Bias: We test if λ_{BlackHost * BlackRater} > 0, indicating that Black participants penalise Black hosts less than White participants do.
Cross-Treatment Comparisons:
While our main hypotheses focus on within-treatment effects, we will explore whether the baseline main effects (e.g., the racial penalty coefficient β_1 differ across the three treatment regressions. This will reveal whether discrimination against minority hosts varies depending on the type of reputation information (quantity vs. quality) available in the market.
Multiple Hypothesis Correction:
We will apply the Benjamini-Hochberg correction to our exploratory hypotheses and interaction effects to control the False Discovery Rate (FDR) at α = 0.10.
Robustness Checks
We will assess the robustness of our main specifications by:
1. Alternative Error Structures: Re-estimating models with different covariance structures, specifically comparing our main clustered standard errors against models with random effects to explicitly model unobserved individual heterogeneity.
2. Functional Form: Verifying that our results hold when using non-linear structural choice models (Conditional Logit or Probit).
3. Budget Stability: Testing whether the estimated penalties are stable across the three budget blocks (Low, Mid, High) by interacting the main effects with the δ_t block dummies.
Pilots
We executed two pilots:
(i) Property Calibration: We initially ran a pre-test with 45 rounds of properties without host information to identify the most comparable pairs of properties. This process allows us to identify the 28 most balanced pairs for the main experiment. These pairs maximise the variance available to be explained by host characteristics.
(ii) Parameter Estimation: We ran a pilot using the fully realised 2x2x2 host/review information to establish a baseline effect of host characteristics on participant choices. We ran this pilot twice varying the incentives: once using the Krupka-Weber method, where participants were rewarded for selecting the property selected by most other participants, and once with participants self-reporting their most preferred property without incentives. This confirmed that: (i) the elicitation methods do not have a statistically significant impact on participant choices, and (ii) the raw marginal differences for demographic traits are in the order of 3 percentage points, informing our simulation-based power analysis.