Minimum detectable effect size for main outcomes (accounting for sample
design and clustering)
Power calculations rely on three empirical benchmarks drawn from prior correspondence studies. First, we use a baseline positive response rate of 13.26% for apprenticeship applications, based on Anne, Bagayoko, Chareyron & L'Horty (2024). Second, we use an alternative and more conservative baseline response rate of 4.5% for apprenticeship applications in a related occupational context, based on Anne, Challe, Chareyron, L'Horty & Noet (2025). Third, we use a hypothesized relative gap in success rates of 15.2% for candidates signaling a visual impairment, based on Challe, Chareyron, L'Horty, Du Parquet & Petit (2023), a correspondence study on housing access.
Combined with the 13.26% baseline, the required combined sample size to detect this effect at 80% power (α = 0.05, two-sided) is 7,940 applications. Combined with the more conservative 4.5% baseline, the required combined sample size is 25,484 applications.
Regardless of which baseline reference rate is used, the required combined sample size (7,940 to 25,484) remains well below our planned sample of 28,000 applications. We are therefore adequately powered to detect the expected effect under either scenario.