Minimum detectable effect size for main outcomes (accounting for sample
design and clustering)
The outcome is mean attribute B investment per matching group, measured in points on a 0 to 100 scale. The matching group is both the unit of randomisation and the unit of analysis, so clustering is fully accounted for by aggregating to the group level before testing.
A pilot session gives a pooled between-group standard deviation of attribute B of 4.0 points, but estimated on only 3 degrees of freedom, so the 95 percent interval for the standard deviation runs from 2.3 to 15.1. We therefore report minimum detectable effects across a range of standard deviations rather than relying on the pilot point estimate.
Because the primary test is nonparametric and the outcome is bounded below at zero, power was simulated rather than taken from a normal-theory formula. The simulation draws the group-level outcome as a normal censored at zero, to reproduce the mass of groups that settle on zero investment, and applies the Mann-Whitney rank-sum test at the 5 percent two-sided level. With 18 groups per arm, the minimum detectable difference at 80 percent power is 4.08 points if the standard deviation is 4, 5.04 points if it is 5, 6.09 points if it is 6, 8.07 points if it is 8, and 10.12 points if it is 10.
The predicted treatment effect is 15 to 16.67 points, and the pilot difference over the analysis window was 14.06 points. Expressed relative to the predicted effect, the design detects effects of roughly 25 to 40 percent of the predicted size for standard deviations between 4 and 6. At the pilot effect size, simulated power exceeds 0.99 for any standard deviation up to 8, and is 0.97 even at a standard deviation of 10.
A t-test applied to the same simulated data gives minimum detectable effects within about 1 percent of the rank test, so the choice between the two does not drive the sample size. A normal-theory t formula would give figures about 5 percent smaller, but that formula assumes an uncensored normal outcome and therefore understates the minimum detectable effect for this design. Simulated size of the rank test is approximately 0.05 under equal variances.