REPUTATION SYSTEMS

Reviewer Honesty Follows a Smile Curve. The Middle Is Where It Breaks.

Newcomers rate candidly because they have not yet learned the room. Then they learn it, inflate for years to keep the peace with whoever they are rating, and only turn honest again once their loyalty shifts from that one person to the community itself.

Based on the research ofMöhlmann, Berente, Son and Lalor, "Inflation in Reputation Systems? Newcomers, Veterans, and Socialization within a Platform Community," Information Systems Research, 2026

Rating honesty forms a smile curve across a reviewer's tenure Candor in ratings tenure on the platform same candor, both ends Newcomer not yet socialized Inflator direct reciprocity to the counterparty Veteran generalized reciprocity to the community peak rating inflation lives here
Rating inflation is not a steady drift upward. It rises through a reviewer's middle years and falls again once loyalty shifts from the person being rated to the community doing the rating.

The first review a newcomer ever posts is, statistically, the most honest one they will write for a long time.

That is the counterintuitive core of a new mixed-method study by Mareike Möhlmann, Nicholas Berente, Yoonseock Son, and John Lalor, published in Information Systems Research. Rating inflation, the well-documented tendency for online reviews to cluster near the top of the scale regardless of actual quality, is usually treated as a one-way ratchet: the longer a reputation system runs, the mushier its scores get. Möhlmann and colleagues found something stranger. Individual reviewers do not simply inflate more the longer they stick around. They move through three archetypal phases. Newcomers, not yet socialized into the community's norms, rate candidly. Inflators, typically the mid-tenure group, absorb a norm of direct reciprocity with the people they are rating and inflate. Veterans, the longest-tenured reviewers, pull back toward candor again, this time driven by a felt commitment to the platform community rather than to any one counterparty. Plotted across a reviewer's lifecycle, inflation traces an inverted U: it rises, peaks, and falls. Honesty does the opposite. It forms a smile curve, high at both ends and lowest in the middle.

The mechanism is reciprocity theory, applied to two different objects. Direct reciprocity is the social contract between two specific parties: you scratch my back, I scratch yours, and if you don't, I remember. Generalized reciprocity has no fixed counterparty; it is a diffuse sense of obligation to the group you belong to, repaid not to the person who helped you but to the next person who needs help. The paper's claim is that reviewers pass through both regimes on their way to becoming veterans, and each regime produces a different rating behavior. A newcomer has no relationship history with the people or businesses they are rating, so there is nothing to reciprocate yet. A mid-tenure reviewer has accumulated exactly enough interactions, sellers who were nice to them, hosts who thanked them personally, drivers who chatted, to feel a specific debt, and inflation is how that debt gets repaid. A veteran has stopped keeping score with any single counterparty and instead measures their reviewing against what the community needs from them: accurate information for the next person. That shift in loyalty, from a person to a population, is what pulls candor back up.

The Newcomer Has Nothing to Repay Yet

Every reputation system has to onboard people who have never rated anything on it before, and most platforms structurally acknowledge that a brand-new participant is a different statistical animal than an established one. Uber's own driver-rating documentation says a new driver starts at a perfect five stars and that the rating "may fluctuate" until they have collected 100 or more ratings, because a handful of early data points are too noisy and too easily skewed to trust as a stable signal. Amazon runs an analogous distinction on the reviewer side of its own marketplace: its Vine program for early product reviews sorts contributors into Silver and Gold tiers, and Silver members are explicitly described as reviewers "still building their track record," with narrower daily product allowances than the Gold tier reviewers who have already demonstrated sustained, trusted reviewing history. Neither of these is the same mechanism the ISR paper documents, one is about the volume of data needed to trust a rating, the other about relationship-based reciprocity, but both point at the same intuition the paper formalizes: tenure is not a cosmetic detail of a reviewer's profile. It is a proxy for how much socialization, and therefore how much accumulated debt to specific counterparties, has had time to build up. A reviewer with zero history has zero debt. That is precisely why the Newcomer Phase in the paper reads as candid by default: there is no one yet to owe a favor to.

The Middle Years, When Reciprocity Takes Over

The Inflator Phase is where direct reciprocity does its damage, and the clearest real-world proof of that mechanism sits in eBay's own feedback history. For years, eBay allowed both buyers and sellers to rate each other, and the platform watched the rate of retaliatory negative feedback climb for four straight years: a buyer would leave a seller an honest negative review, and the seller would retaliate with a negative review of the buyer, whether or not the buyer had done anything wrong. eBay's director of Global Feedback Policy at the time, Brian Burke, described the pattern plainly to Computerworld: the fear of retaliation was suppressing honest buyer feedback and pushing dissatisfied buyers to simply stop shopping on the platform. eBay's fix, effective May 2008, was structural: sellers lost the ability to leave negative or neutral feedback on buyers at all, severing the direct tit-for-tat loop rather than trying to police it case by case.

Airbnb ran the equivalent experiment with data instead of a policy memo. Andrey Fradkin, Elena Grewal, and David Holtz, publishing in Marketing Science, used a large-scale Airbnb field experiment to test what happens when guest and host reviews are hidden from each other until both sides submit, instead of being revealed the moment either party posts. The blind condition reduced retaliation and reciprocation in feedback and produced measurably lower, more candid ratings as a direct result. Left un-hidden, reviews reciprocate each other; hidden, they do not. That is direct reciprocity, isolated and switched off in a controlled setting. It is also, plausibly, the exact engine idling under the Inflator Phase the ISR paper describes: a reviewer mid-tenure has enough of a relationship with hosts, sellers, or drivers that a rating starts to feel like a message sent to a specific person rather than a fact reported to the world. The scale of the resulting drift is visible in the aggregate numbers: Georgios Zervas, Davide Proserpio, and John Byers found that roughly 95 percent of Airbnb listings sit at 4.5 or 5 stars, a ceiling effect that a purely quality-driven distribution should never produce.

Direct reciprocity keeps a relationship alive. Generalized reciprocity keeps a community alive. Only one of those two instincts produces an honest rating.

The Veteran Stops Rating for the Person and Starts Rating for the Platform

The paper's sharpest finding is not that inflation happens, every platform operator already suspects that, it is that it reverses. Veterans, the reviewers with the longest tenure, become more candid again, and the mechanism the authors identify is generalized rather than direct reciprocity: commitment to the platform community as a whole, not to whichever host or seller is in front of them at the moment. A useful real-world echo of that community-facing identity sits in Yelp's Elite Squad, whose tenure structure is coded directly into the badge: reviewers get a Red badge for one to four years of Elite status, Gold for five to nine years, and Black for ten-plus years, explicitly framed by Yelp as recognition of members who are "role models" for the community rather than simply prolific raters. Yelp's own description of what the badge signals, positivity, helpfulness to fellow users, community stewardship, is a fair sketch of the identity shift the ISR paper's Veteran Phase describes: a reviewer whose sense of obligation has migrated from the specific business they are reviewing tonight to the community of future readers who will rely on that review next month. None of this proves Yelp's own reviewers trace the exact inverted-U inflation curve the paper documents on its studied platform. What it does show is that the social architecture the paper's theory requires, a long-tenured reviewer identity built around service to the community rather than to any single counterparty, is not a lab artifact. Platforms build and reward that identity on purpose.

What a Star Rating Actually Tells You Depends on Who Wrote It

Put the two halves together and the operating lesson is specific rather than vague. A five-star rating is not one signal; it is at least three, depending on where its author sits in their own reviewing tenure. A newcomer's five stars probably means the experience really was good, because there is no accumulated debt yet to distort it. A mid-tenure reviewer's five stars is the one to interrogate, because that is exactly the population in whom direct reciprocity has had the most time to take root and the least time to fade. A veteran's five stars is trustworthy again, but for a different reason than the newcomer's: it survives not because the veteran has no relationships to protect, but because their loyalty has outgrown any single one of them.

That reframing matters most for anyone building trust or fraud-scoring systems on top of review data, and it should change how skeptically you read your own platform's numbers. A reputation system that treats every star as fungible, one point is one point, regardless of who cast it or how long they have been casting them, is quietly averaging together three populations with three different honesty profiles and calling the blend "the truth." Suppose, hypothetically, a marketplace ran the numbers and found its mid-tenure reviewer cohort rated 0.3 stars higher, on average, than an otherwise identical newcomer or veteran cohort rating the same sellers; that gap alone would be enough to distort which sellers get promoted in search and which buyers churn after a disappointing purchase. The ISR paper's contribution is not just naming that gap. It is showing that the gap closes on its own, later, if you let reviewers stay long enough to become veterans, which is also the strongest argument for building reputation systems patient enough to let that recovery happen rather than juicing engagement in ways that keep reviewers permanently mid-tenure and permanently inflating.

Sources

  • Mareike Möhlmann, Nicholas Berente, Yoonseock Son, and John Lalor, "Inflation in Reputation Systems? Newcomers, Veterans, and Socialization within a Platform Community," Information Systems Research, 2026 doi.org
  • Andrey Fradkin, Elena Grewal, and David Holtz, "Reciprocity and Unveiling in Two-Sided Reputation Systems: Evidence from an Experiment on Airbnb," Marketing Science, 2021 papers.ssrn.com
  • Georgios Zervas, Davide Proserpio, and John Byers, "A First Look at Online Reputation on Airbnb, Where Every Stay Is Above Average," SSRN papers.ssrn.com
  • Linda Rosencrance, "EBay feedback changes take effect May 19," Computerworld, 2008 computerworld.com
  • "Feedback policy," eBay Help ebay.com
  • "Welcome to the Elite Squad: FAQs," Yelp Official Blog blog.yelp.com
  • "Amazon Vine Program: Is It Worth It? (2026)," goaura.com goaura.com
  • "Understanding driver ratings," Uber Help help.uber.com
← More on the blog