Tally.Press

What a crowdwork wage study actually measured

Task pay and hourly pay are different quantities. The distance between them is where most of the disagreement about crowdwork earnings sits.

In 2018, Kotaro Hara, Abigail Adams, Kristy Milland, Saiph Savage, Chris Callison-Burch and Jeffrey Bigham presented an analysis of earnings on Amazon Mechanical Turk at the CHI conference. Rather than asking workers what they earned, the team used log data captured by a browser plugin that recorded activity as tasks were completed. The dataset covered roughly 3.8 million tasks completed by about 2,700 workers.

The headline figure that circulated afterwards was a median effective wage of about two US dollars an hour, with only a small minority of workers — the paper puts it at around four per cent — reaching the US federal minimum wage of $7.25. Those figures are worth reading slowly, because the word "effective" is doing a great deal of work in that sentence.

The denominator is the argument

A crowdwork task carries a posted price. Dividing that price by the time spent on the task gives one number. The Hara analysis used a wider denominator: it counted time spent searching for tasks, previewing them, reading instructions and abandoning tasks that turned out to be badly specified — activity that produces no payment but occupies the working session. Once unpaid activity enters the denominator, the hourly figure falls, and the size of the fall depends entirely on how much unpaid activity is attributed to paid work.

This is not a dispute about arithmetic. It is a dispute about what counts as work, and it recurs across the literature. The International Labour Organization's 2018 survey of crowdworkers, led by Janine Berg with Marianne Furrer, Ellie Harmon, Uma Rani and M. Six Silberman, reached a similar conclusion by a different route: its respondents across dozens of countries reported spending a substantial share of their platform time on unpaid activity, and the report treats that share as a defining feature of the arrangement rather than as noise.

Where the sample came from

The authors are direct about the limits of their data, and those limits deserve as much attention as the wage figure. The plugin that generated the logs was installed voluntarily. People who install a tracking tool for crowdwork are not a random draw from the population of crowdworkers: they are more likely to be established, higher-volume workers who treat the platform as ongoing work rather than as an occasional activity. Whether that biases the estimate upward or downward is not obvious. Experienced workers are faster and better at avoiding poorly paid tasks, which pushes measured wages up; they also spend more total hours in the system, which changes the weighting.

Browser-recorded time carries its own difficulties. A tab left open during an interruption looks the same to a log as a tab being worked in. The authors apply cutoffs to handle idle periods, and different cutoffs would yield different medians.

Sample composition on this platform has also shifted over time. Djellel Difallah, Elena Filatova and Panagiotis Ipeirotis, tracking the worker population across several years, documented substantial turnover and a demographic mix that changes with recruitment conditions. An estimate drawn from one period does not automatically describe another.

An older question underneath

John Horton and Lydia Chilton, writing in 2010, approached crowdwork pay from the supply side, examining the wage at which workers were willing to accept tasks. Their framing anticipates a problem that persists: platform work is often performed in small increments alongside other activity, so the notion of an hourly wage has to be constructed by the researcher rather than read off a payslip. Different constructions are defensible, and they do not agree.

The useful conclusion is not a number. It is that any single figure for crowdwork pay carries a set of decisions inside it — about unpaid time, about idle time, about who was sampled and when — and that comparing two figures without comparing those decisions produces a confusion rather than a finding.

Sources

  1. Hara, K., Adams, A., Milland, K., Savage, S., Callison-Burch, C., & Bigham, J. P. (2018). A Data-Driven Analysis of Workers' Earnings on Amazon Mechanical Turk. Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems.
  2. Berg, J., Furrer, M., Harmon, E., Rani, U., & Silberman, M. S. (2018). Digital labour platforms and the future of work: Towards decent work in the online world. International Labour Office, Geneva.
  3. Difallah, D., Filatova, E., & Ipeirotis, P. (2018). Demographics and Dynamics of Mechanical Turk Workers. Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining.
  4. Horton, J. J., & Chilton, L. B. (2010). The Labor Economics of Paid Crowdsourcing. Proceedings of the 11th ACM Conference on Electronic Commerce.

Return to the register