PLATFORM GOVERNANCE

A More Accurate Rating Can Provoke More Cheating

The backfire has a rule. When a score sets a price, half-cleaning it can raise gaming, because trust climbs faster than the loophole closes; when a score sets a rank, sharpening it is always safe.

Based on the research ofJoshua S. Gans and Scott Duke Kominers, "Paying for Quality vs. Paying for Rank: When Purifying a Metric Backfires," NBER Working Paper, 2026

Purifying a metric splits two ways, by how it pays out Partly purify a gameable metric SCORE SETS A PRICE (wages, credit terms) Trust climbs faster than gaming exposure falls Gaming can RISE the fix backfires SCORE SETS A RANK (fixed slots, badges) Ranking cancels trust from the incentive Gaming falls, always the fix is safe The reward structure, not the metric, decides whether accuracy helps.
Making a rating more accurate is supposed to reduce cheating; Gans and Kominers show it can do the opposite, and whether it helps or backfires is decided entirely by how the score gets turned into money.

The Two Ways a Metric Pays Out

Every governance team eventually reaches the same reflex. A quality score is being gamed, so you clean it up. You strip out the manipulable variation, tighten the model, close the loophole. Less room to game must mean less gaming. In their 2026 NBER working paper, Joshua Gans and Scott Kominers prove that this reflex is right in exactly one of the two ways a platform can turn a score into a reward, and dangerously wrong in the other.

Their setting is a matching market where agents can inflate a performance metric, and their method is an analytical equilibrium model with closed-form solutions rather than a dataset. The pivot is monetization. When a metric maps continuously into a price, a wage, an interest rate, a credit term, two things move the moment you purify it. The share of the score that is pure noise-and-gaming shrinks, which is the effect you wanted. But the market's confidence in the score also rises, because buyers now believe the number and pay against it. If that trust climbs faster than the gameable variation falls, the payoff to one more unit of gaming goes up, not down, and agents respond by gaming more. A partial clean-up of a priced metric can raise average manipulation. The authors' prescription is uncomfortable for anyone shipping incremental fixes: purify the metric fully or leave it alone, because the halfway house is the trap.

Rank-based allocation is the opposite regime, and it is immune. When a platform hands out a fixed set of positions by ordering, the top ten slots, a certification tier, a badge, all that matters is who sits above whom. The absolute level of trust in the metric cancels out of the incentive, because moving up the ranking is what pays, not the raw number. Sharpen the metric and gaming falls monotonically, every time. There is no confidence channel to reverse the gain. The same act, de-biasing a score, is safe under ranking and potentially self-defeating under pricing.

The reward to gaming is not set by the metric. It is set by how the metric is cashed in.

Why Price Markets Are Fragile

The cleanest priced metric in consumer finance is the credit score, and it shows why the pricing regime is so delicate. A FICO score runs from 300 to 850 and maps almost continuously into the interest rate a borrower is offered. Experian's own rate tables put a 620-score borrower at roughly 7.41% on a 30-year conventional mortgage against about 6.66% for a 760-plus borrower, a spread that works out to around $142 a month on a $350,000 loan. The score is not a label. It is priced to the basis point, and every point of it is money.

That is precisely the environment Gans and Kominers flag as fragile. Around any priced score grows an optimization industry, the credit-repair and tradeline market being the obvious one, dedicated to moving the number without moving the underlying risk. Now imagine a bureau tightening its model to neutralize one such tactic. It closes a loophole, which lowers the exposure to gaming. But a visibly cleaner, better-behaved score also earns more trust from lenders, who then price more aggressively off it. The paper's warning is that if the second effect outruns the first, the partial fix leaves average gaming higher than before, even though the specific loophole is gone. This is a prediction of the model, not a measured outcome in mortgage data, and it should be read that way. But the mechanism is not exotic: it is what happens whenever you make a number more believable and more lucrative to move at the same time.

The managerial tell is whether your accuracy improvement travels with a confidence improvement. A quiet backend fix that no lender notices is safe, because trust does not move. A loudly announced "we've cleaned up the score, you can rely on it now" is the exact combination the model punishes, because it lifts the reward to manipulation across every agent still holding a tactic you did not close.

Rank Is the Safe Regime

Web search is the rank regime at planetary scale, and it behaves the way the theory predicts. Google allocates a fixed, ordered set of positions, and its published spam policies are blunt about the consequence of manipulation: sites that violate the rules "may rank lower in results or not appear in results at all," detected "both through automated systems and, as needed, human review." Those automated systems, the spam-fighting stack Google calls SpamBrain, exist to sharpen the quality signal against an adversarial SEO industry that never stops probing it. Crucially, every time Google refines that signal, the allocation stays purely ordinal, so there is no trust-inflation channel to undo the improvement. Search can sharpen its metric as hard as it likes and only ever squeeze gaming down. That is the paper's rank result running live across billions of queries.

College rankings show the same safety from the audit side. U.S. News hands out ordinal positions, and in 2022 a Columbia mathematics professor, Michael Thaddeus, demonstrated that the university's self-reported figures on class size and faculty credentials were inaccurate. Columbia conceded it had used "outdated and/or incorrect methodologies," was dropped from that year's ranking, and subsequently landed at 18th while moving its reporting to the standardized Common Data Set and hiring external verifiers. Purifying the inputs to a ranking simply reorders the table; it does not create a perverse incentive to game harder, because no continuous reward inflates when the data gets cleaner. Ranked systems still get gamed, of course, schools optimize relentlessly for the formula, but tightening the metric reduces that gaming rather than feeding it.

The Design Question Before the Fix

We are in a purification wave, which is exactly why the distinction matters now. In 2024 the Federal Trade Commission finalized a rule banning fake reviews and testimonials, with civil penalties reaching roughly $51,744 per violation, an economy-wide clean-up of a metric that platforms lean on constantly. The reflex is to treat a de-faked review corpus as an unambiguous win. Gans and Kominers say to ask one question first: how is that review score monetized on your platform? Where reviews set ordering and visibility, a rank-like use, sharpen away and expect gaming to fall. Where reviews feed something price-like, dynamic surge, review-linked commissions, algorithmic payouts, a partial clean-up that also makes buyers trust the stars more carries the backfire risk the model formalizes.

The same authors make the point again in a companion 2026 working paper, "What Does A Grade Mean?", which builds difficulty-adjusted "eigengrades" to strip bias out of academic transcripts. Whether that purification helps or hurts turns on the identical fork: if grades gate ranked, fixed slots, a class rank, an honors cutoff, sharpening them is safe; if they map continuously into scholarship dollars or salary offers, a half-fix inherits the pricing regime's fragility.

So the design variable was never "is our metric accurate." It is "does our metric set a rank or a price." Answer that before you fund the accuracy project, because only then do you know whether making your rating more honest will make your users behave more honestly, or hand the sharpest gamers a better-paid target. Rank-based designs are the safer place to invest in accuracy, and if you can only fix a priced metric partway, the model's counsel is to reconsider whether to fix it at all.

Sources

  • Joshua S. Gans and Scott Duke Kominers, "Paying for Quality vs. Paying for Rank: When Purifying a Metric Backfires," NBER Working Paper 35363, 2026 nber.org
  • Joshua S. Gans and Scott Duke Kominers, "What Does A Grade Mean? Informativeness and Strategic Manipulation of Grading Systems," NBER Working Paper 35183, 2026 nber.org
  • "Average Mortgage Rates by Credit Score," Experian experian.com
  • "Spam Policies for Google Web Search," Google Search Central developers.google.com
  • "Columbia University Admits Submitting Inaccurate Data For Last Year's U.S. News Rankings," Forbes forbes.com
  • "Federal Trade Commission Announces Final Rule Banning Fake Reviews and Testimonials," Federal Trade Commission ftc.gov
← More on the blog