Based on the research ofZhang, Cao, Cui and Zhang, "Evolution of Discrimination on Online Platforms," Management Science, 2026
Audit studies are the workhorse of discrimination research for good reason. Send matched resumes or matched profiles into a market, randomize the names, count the callbacks, and you get a clean causal number nobody can argue with. Marianne Bertrand and Sendhil Mullainathan ran the canonical version two decades ago: identical resumes with White-sounding names got 50 percent more callbacks than the same resumes with Black-sounding names. Benjamin Edelman, Michael Luca, and Dan Svirsky ran the platform-era version on Airbnb: guest applications with distinctively African-American names were 16 percent less likely to be accepted than identical applications with White names. Both are rigorous, both are damning, and both are, by design, a single photograph. A new working paper in Management Science argues that on modern platforms, the photograph is hiding the story.
The Snapshot That Passed
Peibo Zhang, Xinyu Cao, Ruomeng Cui, and Dennis Zhang built their study around an online educational platform where students book one-on-one classes with teachers and can see each teacher's accumulating reviews and booking history. Instead of one audit wave, they tracked real teachers with individual fixed-effects models over their first 12 weeks on the platform. In week one, the gap looks like an ordinary, containable problem: African-American teachers received 30.2 percent fewer bookings than comparable White teachers, and the classes they opened were 3.3 percent less likely to be booked at all. That is a real gap, but it sits in the same neighborhood as plenty of published audit-study findings, the kind of number a platform's trust-and-safety team might read, flag, and file under "monitor."
Then the researchers kept watching the same teachers instead of running a fresh snapshot. Over the following 11 weeks, the initial gap widened by 1,007 percent. Not a new source of bias arriving in week six. Not a different cohort of more prejudiced students showing up later. The same starting gap, left alone, grew elevenfold under its own power. A platform that commissioned a one-time audit in week one, the standard, industry-accepted way of checking for discrimination, would have measured a modest problem and, quite possibly, moved on.
The Machine Behind the Multiplier
The mechanism the authors trace is not "students get more racist over time." It is closer to the opposite, and that is what makes the finding land. Their analysis shows new students are not significantly more likely to skip a Black teacher than a White teacher once the two teachers have minimal or comparable reputation metrics, similar review counts, similar ratings, similar track record. Bias-per-decision, measured against a level playing field, looks close to negligible for that specific comparison.
The problem is that the playing field never levels. Because Black teachers started with 30.2 percent fewer bookings in week one, they accumulate reviews, ratings, and repeat customers more slowly than White teachers from that point forward. Every week that passes, a smaller share of Black teachers has climbed into the "comparable reputation" bracket where the direct bias vanishes, and a larger share of students, new and repeat alike, is choosing between a Black teacher with a thin track record and a White teacher with a thick one. The visible gap widens not because anyone got more biased, but because the population being compared got less comparable. Fewer bookings slow reputation growth; slower reputation growth produces a worse comparison for the next student; a worse comparison produces fewer bookings again. That loop, run for 12 weeks, is what an eleven-times multiplier looks like from the inside.
The bias isn't in what any single new student does. It's in how rarely a Black teacher gets the chance to look identical to a White teacher on the metrics the platform actually shows.
The authors are explicit that this is a platform-design finding, not just a discrimination finding: they show the same amplification mechanism operates regardless of whether the initial gap comes from racial bias at all, which means any early disadvantage, a slow first week, an unlucky first review, a late join date, is vulnerable to being blown up by the exact same reputation math.
Reputation Cuts Both Ways: What Airbnb Already Proved
The unsettling part is that a companion body of research already shows reputation is not merely an amplifier, it is also the single most effective neutralizer anyone has documented for this exact kind of bias. Ruomeng Cui, one of the co-authors of the new study, ran field experiments on Airbnb years earlier with Jun Li and Dennis Zhang and found that guest requests with African-American-sounding names were 19.2 percentage points less likely to be accepted than requests with White-sounding names. But once a guest account picked up a single posted review, positive, non-positive, or even blank, the acceptance-rate gap between White- and Black-sounding names became statistically indistinguishable. One data point of social proof, and the discrimination that self-reported profile information (claims of tidiness, friendliness) couldn't touch simply disappeared.
Read against the new education-platform paper, this is the same lever pointed in the opposite direction. Reputation accumulation is what widened the education-platform gap to 1,007 percent when it was left to run on its own, unsubsidized pace. Reputation accumulation is what erased the Airbnb gap once a guest got even one credible signal fast. The mechanism is identical; the outcome depends entirely on whether the disadvantaged party gets pulled across the "comparable reputation" threshold quickly or slowly. A platform that leaves that threshold to organic, bookings-driven pacing is choosing, by default, the slow version, and the slow version is the one that compounds.
Why This Isn't Just a Tutoring-Platform Story
The pattern generalizes because the underlying plumbing, visibility and trust built from accumulated reviews, ratings, and completed transactions, is common architecture across gig and marketplace platforms, not a quirk of online tutoring. Yanbo Ge, Christopher Knittel, Don MacKenzie, and Stephen Zoepf sent nearly 1,500 controlled ride-hail requests in Seattle and Boston and found African-American-sounding names waited up to 35 percent longer for pickup in Seattle and had their trips canceled by Uber drivers more than twice as often as White-sounding names in Boston. Anikó Hannák, Claudia Wagner, David Garcia, Alan Mislove, Markus Strohmaier, and Christo Wilson pulled 13,500 worker profiles from TaskRabbit and Fiverr and found race and gender significantly correlated with the customer reviews and search rankings that determine which workers get seen at all, the exact reputation infrastructure the Zhang et al. mechanism runs on. Neither study tracked its platform for 12 weeks with individual fixed effects the way the new paper does, so neither can say how much their day-one gaps have already compounded. That is precisely the point: almost none of the discrimination literature is built to answer that question, because almost all of it is audit-shaped.
The Cost of Trusting a Clean Audit
Audit studies remain the right tool for one job: proving discrimination exists at all, with a causal clarity that observational data can't match. What they were never built to do is tell a platform whether its own architecture is quietly turning that discrimination into something much larger by the time anyone checks again. Zhang, Cao, Cui, and Zhang's own framing of the result is blunt: traditional audit studies may significantly underestimate the long-term consequences of early-stage discrimination, because reputation and repeat-customer accumulation function as bias amplifiers on the very platforms doing the auditing.
A platform that ran a one-week test, found a real but modest gap, and moved on has not learned that it treats its Black teachers fairly. It has learned what week one looks like. The 1,007 percent question is what happens by week twelve, and the only way to find out is to keep watching the same people instead of testing a new snapshot.
Sources
- Peibo Zhang, Xinyu Cao, Ruomeng Cui, and Dennis J. Zhang, "Evolution of Discrimination on Online Platforms," Management Science, 2026 doi.org
- Marianne Bertrand and Sendhil Mullainathan, "Are Emily and Greg More Employable Than Lakisha and Jamal? A Field Experiment on Labor Market Discrimination," American Economic Review, 2004 aeaweb.org
- Benjamin G. Edelman, Michael Luca, and Dan Svirsky, "Racial Discrimination in the Sharing Economy: Evidence from a Field Experiment," American Economic Journal: Applied Economics, 2017 aeaweb.org
- Ruomeng Cui, Jun Li, and Dennis Zhang, "Reducing Discrimination with Reviews in the Sharing Economy: Evidence from Field Experiments on Airbnb," Management Science, 2020 papers.ssrn.com
- Yanbo Ge, Christopher R. Knittel, Don MacKenzie, and Stephen Zoepf, "Racial and Gender Discrimination in Transportation Network Companies," NBER Working Paper 22776, 2016 nber.org
- Anikó Hannák, Claudia Wagner, David Garcia, Alan Mislove, Markus Strohmaier, and Christo Wilson, "Bias in Online Freelance Marketplaces: Evidence from TaskRabbit and Fiverr," Proceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing dl.acm.org