Experimental Design
### Video sampling procedure
The videos sampled from this experiment come from four sources:
(A) Videos randomly sampled from the universe of videos posted on TikTok in the US, using TikTok's Research API.
(B) Videos scraped from the recent posts of news organizations, drawing from the list of the top 100 news organizations appearing in the sample of TikTok feeds we collected as part of a previous experiment.
(C) Videos created by participants in our previous "creator" experiment.
(D) Videos scraped from the recent posts of accounts appearing repeatedly in each of the TikTok feeds we collected as part of a previous experiment.
Below, we describe each of these sampling procedures. We then describe how these sources are combined to create participants' experimental feeds.
### A
These are videos scraped from TikTok's Research API across four repeated uses of our daily API quota. We randomly sample from the pool of videos posted 48-96 hours ago in the US. (Videos typically take 48 hours after being posted to appear in the API, so we aren't able to scrape more recently). Half of our API requests sample from the pool of videos with at least 10 views, and half sample from the pool of videos with at least 10k views.
We then scrape the metadata and content of these posts. We exclude posts flagged as ads, posts for which scraping or classification fails due to inaccessible metadata, and posts that our classification flags as being in a language other than English. We also exclude posts whose classification does not fall into one of our four primary content quadrants (lifestyle, produced entertainment, cheap talk, and politics.)
### B
These are videos scraped from the top 100 English-language news organizations appearing in the TikTok feeds we previously collected as part of an experiment. For each organization we scrape their six most recent posts in the previous 7 days. We then scrape and classify them, including according to whether they are a piece of explicitly political news or not.
### C
These are videos we paid participants in our creator experiment $10 to post on their accounts. Each participant was randomly assigned one of our four content quadrants and offered $10 to create a post of this type on their account; these are the posts from people who accepted the offer and sent us the link to the video. As above, we exclude non-English posts and posts not falling into one of our four quadrants.
### D
These are videos scraped from accounts appearing frequently in the feeds of the participants in a previous experiment who provided us with their TikTok viewership histories. For each of these participants, we randomly sample twenty creators who appeared in their feed at least 3 times, replacing creators whose recent profile we fail to scrape. For each of these creators, we scrape their most recent post. These videos are primarily for insertion into the feeds of these previous participants when they now participate in our viewership experiment.
### Constructing participants' feeds
Videos from the aforementioned four sources compose the pool we draw on for the experiment. We then sample posts into participants' feeds stratified in the following way (participants do not have fixed-length feeds, they scroll until 20 minutes have elapsed):
1. For each participant, we independently draw a lifestyle fraction and politics fraction from {0, 0.1, 0.2, 0.3}, with probabilities {0.2, 0.2, 0.4, 0.2}. "Lifestyle" and "politics" means videos from our lifestyle or politics classification quadrants, respectively. Of the remaining videos, 5ppt are non-political news videos posted by news organizations, 1ppt are drawn from the full pool of Source D videos, and the rest are equally split between our cheap-talk and produced-entertainment quadrants. Those are the probabilities for participants who are not part of our previous TikTok-history collection. For participants who are, instead of the 1% drawn from the full pool of Source D videos, they get the 20 videos scraped from creators who previously appeared in their personal feed randomly interleaved into the first 100 posts they see.
2. Within the politics category, 70% of videos are drawn from Source A, 20% from Source B, and 10% from Source C. Nonpolitical news videos are drawn from Source B. Within the other categories, 90% are drawn from Source A and 10% from Source C.