Measuring the Effects of Social Media Without Expensive Surveys
GreenEarth is building models that turn survey experiments into cheap A/B tests
The effects of social media that we might care about most are mostly offline: what people know, believe, and feel. There is certainly information in what users do online but the gold standard has always been surveys — batteries of questions validated by previous psychological research to measure polarization, depression, addiction, and so on.
Off-platform surveys are how essentially every major experimental test of social media effects have worked, at least in part (including the previous work of the GreenEarth team, work from our colleagues, and the big Meta 2020 studies). It makes sense scientifically, but it’s expensive. You have to pay people money to fill out annoying surveys, and we paid an average of $5 per survey last time around — that’s $150,000 for the three waves of our 10,000 user experiment.
Yet people who are depressed or polarized or addicted act differently. They click on different posts and they write different sorts of comments. We know this from previous research. So, in theory, it should be possible to estimate these gold-standard survey results from on-platform data alone.
Part of the GreenEarth project is to test that theory. We’ve built a pro-social feed, now we’re recruiting users, and next we’ll perform a controlled experiment to prove that our feed really is better. We’ll measure “better” using survey results, like our previous experiment, but this time, we’ll also use those results to train models that predict survey answers from on-platform data.

There are serious privacy concerns here which we are taking care to address. First, all training data will come from research participants who have consented to have their data used in this way. Second, we are training models to predict the average survey responses of a group of people — these models cannot be used to predict individual responses, which protects people’s privacy, prevents abuse, and ensures we do not run afoul of laws like GDPR.
There are some deep scientific problems here, notably, it’s not clear how accurately we can actually predict the behavioral and psychological outcomes we care about from on-platform data alone. There must be at least some information that can only be learned offline — but how much? We will be using prediction-powered inference to get robust error bars on all our estimates.
If we succeed, these open models will make it far easier and cheaper for feed designers to estimate the off-platform effects of their algorithms, and provide an easy answer to a question that we’ve heard over and over: suppose I wanted to improve the human outcomes of my feed, what would I even measure?

