Based on the research ofHanna Halaburda, Sarit Markovich, Yaron Yehezkel, "When Does Algorithm Beat Data?," Platform Strategy Symposium, 2026
Hanna Halaburda, Sarit Markovich, and Yaron Yehezkel built a formal model of platform competition in which advantage comes from two distinct sources: how good your algorithm is, and how much data you have accumulated, which in turn drives a data-driven network effect for the incumbent. Presented at the 2026 Platform Strategy Symposium, the paper does something the popular narrative about Big Tech's data moats rarely bothers to do: it derives, rather than assumes, when each source wins. The answer is a threshold, not a verdict. Below it, a superior algorithm overcomes the incumbent's data stock. Above it, the data network effect dominates and the incumbent's lead compounds. Which side of that threshold your market sits on is the single most important strategic fact an operator in a data-intensive platform business can know, and it is knowable, not a matter of faith in whoever has the bigger warehouse.
The Story Everyone Tells Themselves About Data
Walk into any platform strategy conversation about Big Tech and you will hear a version of the same argument: whoever has the most data wins, because more data trains a better model, a better model attracts more users, and more users generate more data. The flywheel spins forever and the incumbent's lead becomes uncatchable. This is also the load-bearing assumption behind a decade of antitrust concern. Charlotte Slaiman of Public Knowledge, testifying before the US Senate Subcommittee on Competition Policy, Antitrust, and Consumer Rights, put the mechanism plainly: dominant platforms merge data across products so that "twice as much data isn't just twice as good for a company, but many times more so," and gatekeepers use that compounding advantage, together with control of the user interface, to pick winners and losers beneath them. It is a coherent story. It is also, according to the formal model at the center of this paper, only sometimes true.
The model does not dispute that data network effects are real, or that they can be decisive. It disputes that they are unconditional. Halaburda, Markovich, and Yehezkel's contribution is to specify exactly where the boundary sits between a world where the entrant's algorithmic edge wins and a world where the incumbent's data stock wins, using analytical comparative statics rather than an empirical estimate from any single market. That is a deliberate choice: a theoretical model built to hold across markets, so a manager can locate their own market on the map instead of importing a conclusion from whichever case study happens to be trending.
Two Sources of Edge, One Threshold
Strip the model to its core and platform advantage has exactly two ingredients. The first is algorithmic quality: how well the underlying model converts a unit of data, or a unit of user interaction, into value. The second is data scale: how much data the incumbent has amassed, and the network effect that scale generates because the platform improves as more users feed it more data. An entrant with a superior algorithm is not competing against the incumbent's user base directly. It is competing against the compounding value of the incumbent's data, run through a comparatively weaker algorithm.
The paper's central result is that these two forces trade off, and the trade-off has a threshold. When the marginal value of additional data is low, an algorithm advantage offsets a large data disadvantage, and the entrant can win outright, sometimes capturing a large share of the market quickly. When the marginal value of additional data stays high, the network effect keeps compounding faster than any algorithm gap can close, and the incumbent's lead is durable. The paper's setting is dynamic rather than static: because an incumbent's existing data also depreciates over time (customer needs shift, past interactions age, what worked last year no longer describes this year's users), even an entrant that starts at zero share can gain ground as the incumbent's stockpile loses relevance faster than the entrant can be out-innovated. Durability of a data advantage, in other words, is itself a variable, not a constant.
DeepSeek Turned the Threshold Into a Live Case
The clearest recent illustration of the algorithm side of that trade-off arrived in January 2025, when the Chinese AI lab DeepSeek released a reasoning model that reporting pegged at roughly $6 million to build, against the hundreds of millions to billions that incumbent labs like OpenAI and Google had been spending, according to CBS News coverage of the market reaction. The response was not a polite research debate. Nvidia shares fell 17% in a single session, the largest one-day market-cap loss in stock market history at roughly $600 billion, and the sell-off spread to chipmakers and even power-infrastructure stocks whose valuations depended on ever-growing AI compute demand. The market was repricing a belief in real time: that the winners in AI would be whoever could spend the most on data and compute. DeepSeek's own technical paper, later published in Nature, reported that reasoning ability could be trained through reinforcement learning rather than the enormous human-labeled data pipelines competitors relied on, and that the resulting model reached performance competitive with the frontier incumbents on verifiable tasks like mathematics and coding.
None of this proves DeepSeek operated with less real infrastructure than critics later alleged, and the training-cost figure itself became contested. But the strategic shock was not about the exact dollar figure. It was that a lab with unambiguously less accumulated data and compute than OpenAI or Google produced output competitive with the frontier, by squeezing more value out of the algorithm side of the ledger. That is precisely the condition the Halaburda-Markovich-Yehezkel model describes: an entrant's algorithmic edge overtaking an incumbent's scale advantage once the marginal value of additional data or additional compute stops buying proportionally more. Whether that condition holds in your market is an empirical question about your specific data's diminishing returns, not a foregone conclusion either way.
The Flywheel Has Diminishing Returns Built In
The reason the algorithm side of the trade-off keeps winning more often than incumbents expect is that data scale rarely behaves like the flywheel it is marketed as. Andreessen Horowitz's Martin Casado and Peter Lauten, examining enterprise AI startups, found that unlike classic economies of scale, the cost of capturing the next unit of useful data tends to rise over time while the value of that unit falls: early data covers the common cases, and everything after covers an increasingly rare, increasingly expensive long tail. In one example they cite, a customer-support chat corpus hit a ceiling where, after roughly 40% of the relevant queries had been collected, additional data added essentially no further coverage. The narrative of "data network effects" as an unstoppable compounding force, they argue, mostly describes a weaker phenomenon, a data scale effect, that in most domains plateaus rather than compounds.
That plateau is exactly the mechanism the formal model formalizes: the marginal value of the incumbent's next data point determines whether the entrant's algorithm can close the gap. A market where relevant data saturates quickly, where user needs are homogeneous, or where the domain shifts fast enough to depreciate old data, is a market where the entrant's algorithmic edge has room to work. A market with a long, genuinely valuable tail of rare, hard-to-synthesize data, the way credit scoring or medical imaging behaves, is a market where the incumbent's stockpile keeps paying off.
More data always wins is not a law of platform competition. It is the case where the threshold happens to favor the incumbent, and the paper tells you how to check which case you are in.
What This Changes for Strategy and for Policy
For an incumbent, the paper's message is not that data moats are fake. It is that a data moat is a claim about a specific threshold in a specific market, and that claim needs verifying rather than assuming. If your domain has a long tail of genuinely scarce, high-value data, hold the line and keep investing in scale. If your domain's marginal data value plateaus quickly, spending on defensibility should shift from hoarding more data toward making the data you already have harder to copy, and toward the switching-cost and multihoming-cost levers that do not depend on the data math at all.
For an entrant, the paper is a targeting tool. It says, in effect, do not waste capital trying to out-collect a company that has been collecting for a decade; instead, find the corner of the market where the marginal value of the incumbent's next data point is already low, and win there with a better algorithm before the incumbent notices the ground has shifted. DeepSeek did not out-data OpenAI or Google. It found the corner of the trade-off where that was not the fight it needed to win.
For regulators and boards debating whether Big Tech's data advantage constitutes a permanent, unassailable moat, the honest answer this model supports is: sometimes, and it depends on conditions that are visible and checkable, not on the sheer size of the incumbent's warehouse. That is a less satisfying headline than "data is destiny." It is also closer to true.
Sources
- Hanna Halaburda, Sarit Markovich, Yaron Yehezkel, "When Does Algorithm Beat Data?," Platform Strategy Symposium, 2026 questromworld.bu.edu
- "The Empty Promise of Data Moats," Andreessen Horowitz a16z.com
- "What is DeepSeek, and why is it causing Nvidia and other stocks to slump?," CBS News cbsnews.com
- DeepSeek-AI et al., "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning," arXiv (published in Nature, 2025) arxiv.org
- Charlotte Slaiman, "How Big Data Fuels Big Tech's Anticompetitive Conduct and Gatekeeping Power," ProMarket (Stigler Center, University of Chicago Booth School of Business) promarket.org
- Robert Wayne Gregory, Ola Henfridsson, Evgeny Kaganer, and Harris Kyriakou, "The Role of Artificial Intelligence and Data Network Effects for Creating User Value," Academy of Management Review doi.org