When Experimental Economics Meets Large Language Models: Evidence-based Tactics

Last registered on July 27, 2026

Pre-Trial

Trial Information

General Information

Title
When Experimental Economics Meets Large Language Models: Evidence-based Tactics
RCT ID
AEARCTR-0019126
Initial registration date
July 25, 2026

Initial registration date is when the trial was registered.

It corresponds to when the registration was submitted to the Registry to be reviewed for publication.

First published
July 27, 2026, 7:07 AM EDT

First published corresponds to when the trial was first made public on the Registry after being reviewed.

Locations

There is information in this trial unavailable to the public. Use the button below to request access.

Request Information

Primary Investigator

Affiliation
Sun Yat-sen University

Other Primary Investigator(s)

PI Affiliation
Tsinghua University
PI Affiliation
Central University of Finance and Economics
PI Affiliation
Tsinghua University
PI Affiliation
Tsinghua University
PI Affiliation
Hong Kong University of Science and Technology

Additional Trial Information

Status
In development
Start date
2026-07-27
End date
2026-09-15
Secondary IDs
Prior work
This trial does not extend or rely on any prior RCTs.
Abstract
Advancements in large language models (LLMs) have sparked a growing interest in measuring and understanding their behavior through experimental economics. However, there is still a lack of established guidelines for designing economic experiments for LLMs. Inspired by principles from experimental economics with insights from LLM research in artificial intelligence, we outline key considerations in the experimental design and implementation stage. We then conduct two sets of LLM experiments to assess how these design choices affect model responses, and two human experiments to examine whether they have comparable effects on human behavior. Our study enhances the design, replicability, and generalizability of LLM experiments, and broadens the scope of experimental economics in the digital age.
External Link(s)

Registration Citation

Citation
Gai, Jianuo et al. 2026. "When Experimental Economics Meets Large Language Models: Evidence-based Tactics." AEA RCT Registry. July 27. https://doi.org/10.1257/rct.19126-1.0
Experimental Details

Interventions

Intervention(s)
Intervention Start Date
2026-07-27
Intervention End Date
2026-09-15

Primary Outcomes

Primary Outcomes (end points)
For Experiment 1, the primary outcome is the Critical Cost Efficiency Index (CCEI). For Experiment 2, the primary outcomes are participants’ decisions in the games.
Primary Outcomes (explanation)

Secondary Outcomes

Secondary Outcomes (end points)
Secondary Outcomes (explanation)

Experimental Design

Experimental Design
Participants in the two experiments will be recruited through Prolific using a representative sample of the U.S. population. They will be redirected to the experimental platform developed using oTree (Chen et al., 2016) and randomly assigned to one treatment arm. All participants will first provide informed consent and then begin the experiment. At the end of the experiment, they will answer several demographic questions and draw a random number to determine whether they receive an additional bonus. The participation fee and any bonus payments will be distributed through Prolific within one week.

Experiment 1:
Experiment 1 involves budgetary decision tasks. Participants will complete 25 rounds of decision task in the risk preference domain (Choi et al., 2007, 2014) and 25 rounds in the social preference domain (Andreoni and Miller, 2002). The order of the two tasks will be randomized. In the risk preference domain, the participant allocates 100 points to two accounts, where the exchange rates between points and payoffs differ. She receives the payoff from one random account, selected with equal probability. For the social preference, the participant allocates 100 points between herself and another randomly matched agent. These points are converted into payoffs with distinct exchange rates for the two agents. Budget lines vary across rounds for both tasks. We will randomly select, on average, 1 out of every 30 participants to receive additional bonuses. For each of the selected participants, we will randomly select one round across total 50 rounds to realize her additional bonus.

Experiment 1 consists of three arms, examining whether dialogue structure affects economic rationality. In the first arm (Single-turn dialogue), all 25 rounds of decision are presented on the same page. In the second arm (Multi-turn dialogue), each page presents a new round of the task while retaining all previous questions and responses. In the third arm (Repeated single-turn dialogue), each page presents a new round of the task without retaining any previous questions or responses.

Experiment 2:
Experiment 2 involves five decision scenarios drawn from four tasks: the dictator game, the ultimatum game, the public goods game, and the bomb risk elicitation task. The order of the decision scenarios will be randomized. In each game, participants are required to make one decision, with the following exception—in the ultimatum game, the participants separately play the role of proposer, who proposes an allocation, and responder, who determines the minimum proposal they are willing to accept. Detailed descriptions of these tasks can be found in Mei et al. (2024). To avoid learning effects and simplify the analysis, we focus solely on the first round of decisions in each scenario. We will randomly select, on average, 1 out of every 30 participants to receive additional bonuses. For each of the selected participants, we will randomly select one game to realize her additional bonus.

Experiment 2 consists of two arms, examining whether answer format affects economic preferences. In the first arm (Open-ended answer format), participants may choose any value within the feasible range. In the second arm (Multiple-choice answer format), participants can select only one of the provided options. Given that the feasible choice sets in all scenarios except the public goods game range from 0 to 100, we discretize this range into 21 options in increments of 5. In the public goods game, where the feasible choice set ranges from 0 to 20, we retain 21 options and use increments of 1.
Experimental Design Details
Not available
Randomization Method
Randomization algorithm embedded in the otree.
Randomization Unit
Individual.
Was the treatment clustered?
No

Experiment Characteristics

Sample size: planned number of clusters
300 individuals for Experiment 1 and 200 individuals for Experiment 2.
Sample size: planned number of observations
300 individuals for Experiment 1 and 200 individuals for Experiment 2.
Sample size (or number of clusters) by treatment arms
100 individuals in each treatment arm for both experiments.
Minimum detectable effect size for main outcomes (accounting for sample design and clustering)
IRB

Institutional Review Boards (IRBs)

IRB Name
Scientific Review Committee at the Central University of Finance and Economics
IRB Approval Date
2026-07-21
IRB Approval Number
IRB20260708001