Primary Outcomes (explanation)
We pre-specify six families of outcomes outlined below, each of which will also include a summary index:
1. Policymaker Index — pooled across our four summary indices at the policymaker level
2. Evidence Use Capacity Index (policymaker-level):
• Frequency of searching for or reviewing evidence: Response to self-report of how often in the last three months the policymaker has searched for or reviewed evidence. The same question format is used for the next six questions.
• Frequency of attending research forums
• Frequency of collecting data
• Frequency of running an evaluation
• Frequency of using evidence in policy decisions
• Frequency of citing data or evidence
• Frequency of sharing evidence
• Self-efficacy: Responses to the question “I find it difficult to use or analyze data/evidence.” (Not at all relevant,..., Extremely relevant)
• Quiz performance: Share correct among the seven incentivized quiz questions randomly shown to each respondent.
3. Perceived Organizational Support for Evidence Use Index (policymaker-level):
• Leadership support for evidence use: Response to the statement “Leaders in my unit support the use of data and evidence.” (No not at all,..., Yes very much)
• Organizational training and resources for evidence use: Response to the statement “My unit provides training in using data / resources for evidence (e.g., journal subscriptions, software).” (No not at all,..., Yes very much)
• Organizational processes for evidence review and evaluation: Response to the statement “My unit has a documented process for reviewing evidence / evaluating policies and programs.” (No not at all,..., Yes very much)
• Predicted colleagues’ preferences for evidence use: Incentivized predictions (Krupka-Weber) of how the typical policymaker in the respondent’s unit values quantitative analysis.
4. Preferences Index (policymaker-level):
• Willingness to pay for an evidence tool: Switch point for the choice between a donation to public funds versus an annual Elicit subscription (or similar product), which uses AI to summarize academic research.
• Willingness to engage with a research brief: Binary choice to receive an AI-in-government research brief.
• Willingness to share a research brief: Binary outcome indicating whether policymakers were willing to enter a colleague’s email address to share the AI-in-government brief.
• Conjoint preferences: Choices between evidence quality versus a benchmark attribute, political popularity. In our index we aggregate our evidence attributes and classify preferences as shifting in the direction of being evidence-informed when treated policymakers are less likely to choose before-after comparisons and more likely to choose an RCT.
5. Evidence to Action Index (policymaker-level):
• Ability to measure and collect data on policy outcomes: Lab staff self-reported assessment and indicator for average availability across LLM interviews.
• Change in the policy outcomes identified at baseline: Policymakers specify priority outcomes for their unit at baseline. Improvement in these outcomes are measured over the course of the study, using admin data (whenever available) or self-reports. At baseline, policymakers provide the minimum improvement in these priority outcomes over a year that they would consider a success. Improvements are scaled relative to this benchmark.
• Change in which interventions/programs to pursue based on evidence: At baseline, policymakers are asked to indicate the interventions/programs under consideration in their unit. Policymakers are asked to explain changes in their ranking at midline and
endline. Changes based on data or evidence are coded as a binary outcome.
• Tractability of policy outcomes: Expert evaluation (by Lab staff, researchers, or trained AI model) of the tractability of policy outcomes and success thresholds specified in LLM interviews.
• Tractability of intervention: Expert evaluation (by Lab staff, researchers, or trained AI model) of the tractability and potential impact of interventions specified in LLM interviews.
6. Evidence Adoption Index (unit-level):
• Document analysis (extensive): Proportion of documents and public communications produced by the unit (and meeting notes if available) that reference research evidence.
• Document analysis (intensive): Intensity of references to research evidence in documents and public communications produced by the unit (and meeting notes if available).
• Commissioning of evaluations: Self-reports in survey validated by tracking of policy documents.