Back to History

Fields Changed

Registration

Field Before After
Trial End Date December 31, 2026 December 01, 2026
Last Published April 27, 2026 11:02 AM September 02, 2026 10:33 PM
Study Withdrawn No
Intervention Completion Date June 01, 2026
Data Collection Complete Yes
Final Sample Size: Number of Clusters (Unit of Randomization) The experiment included 282,499 platform-device-shipment randomization units: 73,204 in Phase I and 209,295 in Phase II. The experiment was not cluster-randomized; these platform-device-shipment combinations were the units actually assigned to treatment.
Was attrition correlated with treatment status? No
Final Sample Size: Total Number of Observations The platform initially recorded 1,242,309 raw shipment-tracking query-arm rows. After collapsing duplicate logging records for the same platform-device-shipment-session-arm combination, the final data contained 838,158 distinct query-session exposure records: 229,669 in Phase I and 608,489 in Phase II. These exposure records were generated by 282,499 platform-device-shipment randomization units. Repeated query sessions are post-assignment behavioral observations rather than additional randomization units.
Final Sample Size (or Number of Clusters) by Treatment Arms The final sample contained 282,499 platform-device-shipment randomization units. Because assignment was hierarchical, the 21 treatment arms were not designed to have equal sample sizes. The traditional prediction model with delivery-time-only disclosure included 18,760 units at one-day precision, 17,911 at six-hour precision, 17,772 at one-hour precision, and 18,844 at thirty-minute precision. The AI maximum-probability model with full-route disclosure included 10,021 units at one-day precision, 9,391 at six-hour precision, 9,171 at one-hour precision, and 9,983 at thirty-minute precision. The AI maximum-probability model with delivery-time-only disclosure included 9,378 units at one-day precision, 9,114 at six-hour precision, 8,712 at one-hour precision, and 9,422 at thirty-minute precision. The AI fastest-time model with full-route disclosure included 9,406 units at one-day precision, 9,011 at six-hour precision, 8,761 at one-hour precision, and 9,543 at thirty-minute precision. The AI fastest-time model with delivery-time-only disclosure included 9,283 units at one-day precision, 9,023 at six-hour precision, 8,856 at one-hour precision, and 9,476 at thirty-minute precision. The no-prediction condition included 60,661 units.The final sample contained 282,499 platform-device-shipment randomization units. Because assignment was hierarchical, the 21 treatment arms were not designed to have equal sample sizes. The traditional prediction model with delivery-time-only disclosure included 18,760 units at one-day precision, 17,911 at six-hour precision, 17,772 at one-hour precision, and 18,844 at thirty-minute precision. The AI maximum-probability model with full-route disclosure included 10,021 units at one-day precision, 9,391 at six-hour precision, 9,171 at one-hour precision, and 9,983 at thirty-minute precision. The AI maximum-probability model with delivery-time-only disclosure included 9,378 units at one-day precision, 9,114 at six-hour precision, 8,712 at one-hour precision, and 9,422 at thirty-minute precision. The AI fastest-time model with full-route disclosure included 9,406 units at one-day precision, 9,011 at six-hour precision, 8,761 at one-hour precision, and 9,543 at thirty-minute precision. The AI fastest-time model with delivery-time-only disclosure included 9,283 units at one-day precision, 9,023 at six-hour precision, 8,856 at one-hour precision, and 9,476 at thirty-minute precision. The no-prediction condition included 60,661 units.
Is there a restricted access data set available on request? No
Program Files No
Data Collection Completion Date June 10, 2026
Is data available for public use? No
Intervention (Public) Participants will be exposed to different versions of a package-tracking interface on a digital logistics platform. The intervention consists of randomized variation in how predicted delivery information is presented to users. The study varies three dimensions of the interface: 1. Prediction model type: Users are assigned to different prediction systems, including a traditional statistical prediction model, a higher-accuracy AI-based model, a faster but lower-cost AI model, or a baseline non-model estimate. 2. Information display structure Users are randomly assigned to either view only the final predicted delivery time or view additional intermediate shipment checkpoint information during the delivery process. 3. Time precision of prediction The estimated delivery time is displayed at different levels of precision, including day-level, 6-hour window, 1-hour window, or 30-minute window. After the initial randomized assignment, users may choose to adjust the displayed precision level according to their own preference. These choices are recorded as part of the study. The intervention is embedded in the platform’s normal shipment-tracking service and does not affect the actual delivery process. The intervention was embedded in the shipment-tracking interface of a large digital logistics platform in China. It randomized how predicted logistics information was generated and disclosed when a device queried a particular shipment. The intervention affected only the information shown to users and did not alter the shipment, delivery process, or courier operations. The experiment varied three dimensions: 1. Prediction technology and availability. Shipment queries were assigned to a traditional prediction model, an AI maximum-probability prediction rule, an AI fastest-time prediction rule, or a no-prediction condition. 2. Route-information disclosure. Within the two AI conditions, the interface displayed either only the predicted final delivery time or the predicted final delivery time together with predicted intermediate logistics nodes. The traditional model did not generate intermediate-node predictions. 3. Display precision. Predictions were displayed using one of four randomly assigned time intervals: one day, six hours, one hour, or thirty minutes. These dimensions produced 21 experimental arms: four traditional-model arms, sixteen AI arms defined by two AI prediction rules, two route-disclosure conditions, and four precision levels, and one no-prediction arm. Assignment was hierarchical rather than uniform across the 21 final arms. The intervention was conducted in two phases. Phase I ran from May 3 through May 12, 2026. In this phase, the randomized precision determined the default display, but users could change the precision before confirming their choice. Phase II ran from May 13 through June 1, 2026. In this phase, users could not change the randomly assigned display precision. The remaining treatment dimensions and the underlying shipment-tracking service were unchanged across phases. Treatment was assigned when a platform-device-shipment combination first entered the experiment and was subsequently retained for later queries of the same shipment by the same device.
Experimental Design (Public) This study is a large-scale randomized field experiment embedded in the shipment-tracking interface of a digital logistics platform. The experiment uses a multi-arm factorial design that varies three dimensions of the user interface and prediction system: Prediction model type (4 arms) Users are assigned to one of four prediction approaches: a traditional model, a high-accuracy AI model, a fast AI model, or a baseline non-model estimate. Intermediate shipment information (2 arms) Users are randomly assigned to either view intermediate shipment checkpoint information or only the final estimated delivery time. Prediction time precision (4 arms) Estimated delivery times are displayed at one of four levels of precision: day-level, 6-hour window, 1-hour window, or 30-minute window. These three dimensions generate a factorial treatment structure with multiple treatment combinations. Eligible users are randomly assigned at first exposure to one treatment condition by the platform’s backend experimental system. After the initial assignment, users may choose to adjust the displayed precision level according to their own preferences. This feature allows the study to observe both the causal effects of default information presentation and users’ endogenous preferences for information granularity. The intervention affects only the presentation of predictive information and does not alter the actual logistics or delivery process. This study was a large-scale randomized field experiment embedded in the shipment-tracking interface of a digital logistics platform. The randomization unit was a platform-device-shipment combination. When a device first queried a particular shipment during the experiment, the platform’s backend system used a computer-generated pseudo-random procedure to assign that combination to an experimental condition. Later queries of the same shipment by the same device retained the initial assignment. The experiment used a 21-arm factorial structure. Twenty prediction arms varied three dimensions: prediction technology, display precision, and, where applicable, route-information disclosure. The prediction technologies were a traditional model, an AI maximum-probability rule, and an AI fastest-time rule. Display precision was randomized among one day, six hours, one hour, and thirty minutes. Within the two AI conditions, route disclosure was additionally randomized between displaying only the final delivery-time prediction and displaying the final prediction together with predicted intermediate logistics nodes. Because the traditional model did not generate intermediate-node forecasts, route disclosure was not randomized in that condition. A separate twenty-first arm displayed no prediction. Assignment was hierarchical rather than uniform across the 21 final arms. The backend first assigned the broad prediction condition or technology, then randomized route disclosure within the AI conditions, and finally randomized display precision within each applicable model-by-display cell. The experiment was conducted in two phases. Phase I ran from May 3 through May 12, 2026. Random assignment determined the default display precision, but users could change it before confirmation. Phase II ran from May 13 through June 1, 2026. Users could not change the randomly assigned precision. Thus, Phase I identifies the effect of a randomized default and provides information about endogenous precision choice, while Phase II identifies the effect of fixed display precision. The intervention changed only the prediction and information displayed in the tracking interface; it did not affect shipment handling or delivery. User queries and feedback were recorded during the experimental period. Shipment-trajectory records were retained through June 10, 2026, to recover realized node and delivery times for shipments that had not yet arrived when the intervention ended.
Randomization Method Randomization is performed automatically by the platform’s backend computer system at the time of a user’s first shipment-tracking interaction during the study period. Upon first exposure, each eligible user is assigned by a computer-generated pseudo-random algorithm to one treatment cell in a factorial design defined by three dimensions: prediction model type, intermediate-node display, and time-precision level. The assignment procedure is implemented server-side and does not involve manual intervention, public lottery, or any physical randomization device. The randomization algorithm uses uniform random assignment across all treatment arms and is logged in the platform’s internal experimental framework to ensure reproducibility and auditability. Randomization was conducted automatically by the platform’s backend computer system when an eligible platform-device-shipment combination first entered the experiment. A computer-generated pseudo-random algorithm assigned the prediction condition and, conditional on that assignment, the applicable prediction technology, route-disclosure condition, and display precision. The procedure was hierarchical and therefore did not assign equal probabilities to all 21 final arms. Assignment was implemented without manual intervention and recorded in the platform’s experimental logs.
Randomization Unit The primary unit of randomization is the individual user. Each user is randomly assigned to one treatment combination at the time of first exposure to the shipment-tracking interface during the experimental period, and this initial assignment remains the default treatment condition for that user throughout the observation window. Because multiple shipment queries may be observed for the same user, outcome data will be recorded at both the user level and the shipment-query level, but treatment assignment occurs at the individual-user level. There is no higher-level cluster randomization (e.g., by city, courier company, or shipment route) in the baseline design. The randomization unit was a platform-device-shipment combination, identified by platform, device identifier, and shipment waybill. Treatment was assigned at the first experimental query for that combination and retained for its subsequent queries. Different shipments queried by the same device constituted separate randomization units and could receive different assignments. All treatment dimensions were randomized at this same level. There was no higher-level randomization by user account, city, courier company, route, or other cluster.
Planned Number of Clusters Not clustered and 1,000,000 individual users (randomization units). The experiment was not cluster-randomized. Approximately 1,000,000 platform-device-shipment combinations were planned as randomization units.
Planned Number of Observations Approximately 1,000,000 user-level shipment-tracking sessions. Approximately 1,400,000 shipment-tracking query-session records during the experimental period, while multiple query-session records could be observed for the same platform-device-shipment randomization unit.
Sample size (or number of clusters) by treatment arms 21 treatment arms with approximately 47,620 observations in each arm under equal random assignment. Assignment was hierarchical, so the 21 final arms were not designed to have equal sample sizes. Of approximately 1,000,000 planned randomization units, the four broad conditions were expected to receive approximately 250,000 units each: No prediction: approximately 250,000 units in one arm. Traditional prediction model: approximately 250,000 units, divided equally across four precision arms, or approximately 62,500 units per arm. AI maximum-probability rule: approximately 250,000 units, divided equally across two route-disclosure conditions and four precision levels, or approximately 31,250 units per arm across eight arms. AI fastest-time rule: approximately 250,000 units, divided equally across two route-disclosure conditions and four precision levels, or approximately 31,250 units per arm across eight arms. Actual realized sample sizes could differ because of platform traffic, logging availability, and the timing of eligible shipment queries.
Secondary Outcomes (End Points) 1. Platform economic outcomes: Estimated advertising impressions per shipment Estimated advertising revenue per tracking session Subsequent platform usage within the observation window 2. Feedback detail outcomes: Specific reasons selected for negative evaluations User response conditional on realized prediction The primary outcomes fall into three categories: 1. User feedback. For each eligible shipment-tracking exposure, we measure whether the user records any feedback, positive feedback, or negative feedback. We also construct an indicator for negative feedback accompanied by a specific reason explicitly related to the prediction. Feedback outcomes are matched to the relevant query using the platform, application session, and shipment identifier. 2. Subsequent query behavior. At the platform-device-shipment level, we measure whether the shipment is queried again within one hour, six hours, and twenty-four hours after its first experimental query; the number of additional queries within twenty-four hours; the logarithm of one plus the number of additional queries; and the elapsed time until the next query. 3. Precision choice in Phase I. Among users who interact with the precision-selection interface, we measure whether the final precision differs from the randomized default, whether the user chooses a finer or coarser precision, the final selected precision, and the time spent making the selection where click-level timing is available. The primary causal analysis of fixed display precision uses Phase II, in which users could not change the randomized precision. Phase I is used to estimate the effect of randomized default precision and to study active precision choice.
Back to top