Based on the research ofLi and Zhang, "Strategic Participation on Tokenized Platforms: Balancing Investment and Labor Intensities," MIS Quarterly, 2026
One Wallet, One Clock, Two Roles
A gig driver decides how many hours to work. A retail investor decides how much capital to put at risk. Almost nowhere outside tokenized platforms does the same person have to solve both problems at once, with the same constrained resources, inside the same market. On a blockchain-based platform, a GameFi title, a staking-and-questing protocol, a play-to-earn game, a participant can buy and hold tokens or NFTs as an investor, complete tasks and play to earn token rewards as a laborer, and do both simultaneously with the hours and the capital they have on hand. Tianyi Li and Xiaoquan (Michael) Zhang's paper in *MIS Quarterly*, "Strategic Participation on Tokenized Platforms: Balancing Investment and Labor Intensities," starts from exactly this observation: participants on these platforms "can simultaneously take multiple roles, such as user, investor, and laborer, and draw income from the last two roles," and unlike traditional markets that "typically prioritize one means of profitable participation," tokenized-platform participants have to allocate effort across both roles to maximize what they take home.
That framing sounds like a modest technical point, but it reorganizes the whole decision. In a traditional market, more capital invested and more hours worked are separate optimization problems that happen to share a bank account. On a tokenized platform, they compete directly: an hour on in-game tasks is an hour not spent researching which token to buy, and a dollar staked into an NFT is a dollar that cannot cover the gas fees or in-game purchases labor requires. Li and Zhang build a decision framework for this joint problem, distinguishing an individual's optimal strategy from the platform-average participant's, and splitting the decision into two subproblems: one that treats the platform's future state as given and builds strategy off metrics derived from Monte Carlo ensembles of likely price paths, and one that treats the participant's own actions as feeding back into the token economy, formalized as a Markov decision process and solved with reinforcement learning.
The Fixed-Ratio Instinct
Faced with a two-armed allocation problem under uncertainty, most people reach for a rule. Put some fraction of your budget into tokens, spend the rest of your time and money playing or completing tasks, and keep that ratio roughly stable regardless of what the token price is doing this week. It feels disciplined, the way a 60/40 stock-bond split feels disciplined to a retail investor: a pre-committed rule that resists panic and overtrading.
Suppose, purely as an illustrative hypothetical, a scholar on a play-to-earn platform decides up front to put 70 percent of their available capital into buying in-game assets and reserve 30 percent as a cash buffer, while committing to a fixed three hours a day of play regardless of how token rewards are trending. That is a fixed-ratio strategy in the sense Li and Zhang's metric-based approach describes: a rule built from historical patterns and held constant, rather than a rule that moves with the platform's evolving state.
The trouble is that a tokenized platform's state is not stationary. Token supply inflates as more participants grind for rewards; token price falls as new tokens hit the market faster than new demand arrives; the value of an hour of labor and the value of a dollar of stake can each collapse within weeks, sometimes days. A ratio chosen on day one has no mechanism for noticing that the ground underneath it moved. Li and Zhang's paper exists specifically because that gap is not academic: they test metric-based strategies built on Monte Carlo ensembles against reinforcement-learning strategies that explicitly model the participant's actions as influencing platform state, using real historical token price series, and the RL approach wins.
What the Model Actually Compares
The first subproblem ignores the fact that one participant's actions move the platform at all, a reasonable simplification when a platform has many small participants, and instead derives a strategy from metrics computed over Monte Carlo ensembles: many simulated future paths of the token economy, summarized into statistics a participant can act on. This metric-based strategy is a more sophisticated cousin of a fixed ratio, informed by simulation rather than instinct, but still built without any feedback loop between the participant's own choices and what happens next on the platform.
The second subproblem drops that simplification. It treats the participant's allocation between investing and laboring as an action in a Markov decision process, where today's allocation changes tomorrow's platform state, which changes which allocation is optimal the day after. Solved through reinforcement learning, this strategy adapts continuously: when conditions signal an oversupplied, depreciating token, it shifts weight away from illiquid investment and toward the roles that still pay out; when conditions favor holding, it shifts the other way. Li and Zhang compared the two strategy families using historical token price series, and the RL-based approach outperformed the metric-based one. The mechanism is not that reinforcement learning is a fancier label for the same intuition; it is that a strategy treating its own actions as part of the system it navigates can out-earn one that treats the system as something happening to it.
A rule chosen on day one has no way of noticing that the ground underneath it has moved.
Axie Infinity and STEPN Paid for Guessing
Two real play-to-earn platforms show what it costs to run the fixed-ratio instinct against a platform whose state is actively shifting, even though neither is the dataset in Li and Zhang's study. Axie Infinity, the Vietnam-based blockchain game built by Sky Mavis, built its entire growth model on the investor-laborer split the paper describes. Players bought teams of NFT creatures called Axies, the investment, and then battled for "Smooth Love Potion" (SLP) token rewards, the labor. Because a starter team could cost hundreds of dollars, "scholarship" guilds emerged to split the two roles across two people: an owner supplied the capital, a scholar supplied the hours, and a fixed formula split the proceeds. Yield Guild Games, one of the largest such guilds, reported that in February 2022 its scholars kept 70 percent of what they farmed, with 20 percent going to the scholarship manager and 10 percent to the Guild's treasury, a three-way ratio set by agreement, not by the token's trajectory. By that same month YGG had 20,700 active Axie scholars, up roughly 8,500 percent from 241 a year earlier, with scholars and managers reporting their household incomes transformed almost overnight, and monthly winnings often cited around $200 against households earning less than $400 a month.
That 70/20/10 split is a fixed-ratio allocation institutionalized as a contract: the owner's capital, the manager's coordination, and the scholar's hours were locked into a set share for the life of the agreement, regardless of what happened to the token underneath it. It held exactly the same proportions while SLP rocketed from 3.5 cents to 36.5 cents in roughly a week in spring 2021, and it held the same proportions as SLP sank more than 95 percent from its July 2021 peak of 40 cents to about 1.8 cents by March 2022. A ratio fixed by contract has no mechanism for reallocating anyone's effort as the platform's own token supply and price state deteriorates; the contract itself is the fixed ratio, frozen at signing, indifferent to the Markov process running underneath it.
STEPN, a Solana-based "move-to-earn" app, shows the same pattern on the individual-participant side rather than the contractual side. Users bought NFT sneakers (investment) to earn Green Satoshi Token, GST, for walking or running (labor), with a second token, GMT, layered in as a governance asset bought back and burned out of platform profit, STEPN reported $26.8 million in Q1 2022 marketplace and royalty revenue dedicated to that buyback. Both tokens are a case study in how fast platform state can turn: GMT rose from about one cent on March 9, 2022 to a record $3.45 on April 19, a documented 34,000 percent move in 41 days, while GST hit its own all-time high of $9.03 on April 28, 2022, only to lose more than 45 percent of its value within a day and close that year down more than 99 percent. A participant who bought an NFT sneaker near the peak, committing to a fixed daily walking routine to earn back that stake, had no lever inside a fixed plan to notice that the token paying for the sneaker was about to evaporate, and no mechanism to shift weight toward cashing out the investment, selling the sneaker, or reducing labor hours before the reward stopped covering the cost of showing up.
The Design Lesson for Platforms and Participants
The broader claim in Li and Zhang's paper is not "reinforcement learning beats heuristics," a claim true of enough domains to be unremarkable on its own. It is that on a tokenized platform specifically, the investor role and the laborer role are not two separate decisions that happen to share a participant; they are one resource-allocation problem, constrained by the same capital and the same hours, embedded in a token economy whose state changes partly because of what participants, in aggregate, choose to do. A metric-based strategy is a real improvement over no strategy at all, but it still treats the platform as a weather system to be forecast rather than a system the participant is inside of. The strategy that explicitly models the feedback loop between individual action and platform state is the one that wins when tested against historical token price series.
For a participant, the implication is that "pick a sensible split and hold it" is exactly the instinct a platform in flux will exploit. For a platform designer, the implication is sharper: a user interface that shows only today's reward rate and today's token price, with no visibility into supply trends or reward decay, nudges every participant toward the fixed-ratio mistake by default. Axie Infinity's and STEPN's crashes did not punish anyone for failing to predict the future. They punished a structure that gave almost nobody a way to notice the present was already changing.
Sources
- Li, Tianyi, and Xiaoquan (Michael) Zhang, "Strategic Participation on Tokenized Platforms: Balancing Investment and Labor Intensities," MIS Quarterly, 2026 doi.org
- "'Life-changing' or scam? Axie Infinity helps Philippines' poor earn," AFP via Gulf News, February 15, 2022 gulfnews.com
- "Solana's STEPN hits record high as GMT price skyrockets 34,000% in over a month," Cointelegraph, April 19, 2022 cointelegraph.com
- "Green Satoshi Token Price Prediction | Is GST a Good Investment?," Capital.com Research Team capital.com
- "Yield Guild Games hits 20K Axie Infinity P2E scholarship milestone," FXStreet / Cointelegraph, March 8, 2022 fxstreet.com
- "Announcement: STEPN 1st Quarterly GMT buyback & burn," STEPN Official, April 1, 2022 stepnofficial.medium.com