Based on the research ofSekiya, Otani, Komatsu, Fujii, Ozeki, and Noda, "Designing Recommendation Exposure and Favorite Lists: A Field Experiment in a Spot-Work Platform," arXiv preprint, 2026
Popularity Measures Yesterday's Clicks, Not Today's Vacancy
Timee is Japan's largest spot-work marketplace: an app where workers pick up shifts as short as one hour at convenience stores, restaurants, and hotels, get paid quickly, and never fill out a resume. The company went public on the Tokyo Stock Exchange in July 2024 at a roughly $1 billion valuation, had 7 million registered users at listing, and has since built a worker database of 13.4 million. It exists because Japan's labor market is genuinely short-handed: Recruit Works Institute projects an 11-million-worker shortfall by 2040, and the country's spot-work segment was on track to roughly quintuple to 100 billion yen by fiscal 2028, according to figures disclosed by rival platform Dip. This is not a niche feature. It is infrastructure for an economy that cannot staff itself through conventional hiring.
Timee's core mechanic is the favorite list: a worker favorites a job template, such as "convenience store cashier, evenings, near Shibuya station," and the recommender notifies them whenever an employer posts a new shift matching that template. The natural way to rank which templates get pushed hardest is to predict favoriting, the same logic every recommender uses, from streaming to shopping. And it fails for exactly the reason it usually succeeds elsewhere: it rewards what people have already liked, not what is currently available.
A team of six researchers studying Timee's production system, Sekiya, Otani, Komatsu, Fujii, Ozeki, and Noda, found that maximizing predicted favoriting generates what they call misdirected concentration. Recommendations pile onto templates that are popular precisely because many workers already favorited them, chasing a fixed or shrinking pool of shifts those employers actually post. Meanwhile, templates tied to businesses with real, unmet staffing need, the shifts most likely to convert into an actual job, get comparatively little exposure because they haven't accumulated the favoriting history that would earn them algorithmic attention. Popularity is a lagging indicator of past interest. It says nothing about whether an opening exists right now, and in a marketplace where jobs expire within hours or days, that gap is the whole ballgame.
The Fix: Make Exposure Follow Capacity, Not History
The researchers' answer is thresholded eligibility control, TEC for short: an exposure-control mechanism that reallocates which templates get pushed based on posting activity and unfilled capacity rather than historical favoriting alone. Templates that show signs of genuine, currently unmet demand become eligible for a bigger share of recommendation slots; templates that are already saturated with far more interested workers than openings get throttled back, even if they remain popular by every conventional metric. The mechanism is explicitly designed to be fully parallelizable, meaning it can run at the scale of a national platform processing millions of daily recommendation decisions without a centralized bottleneck.
The results are the part that should make any marketplace operator sit up. In simulations calibrated to Timee's actual data, TEC raised the per-round job-finding rate from 57.6% to 70.0%, a 12.4-percentage-point jump achieved purely by changing which templates get shown to whom, with no change to the underlying supply of jobs or workers. That is the headline number in the exhibit above, and it is worth sitting with: the same jobs, the same workers, a different exposure rule, and roughly one in eight additional job searches now ends in a match.
Simulation gains are cheap to produce and easy to overstate, so the paper's more important contribution is the field experiment. Timee ran a prefecture-level randomized rollout of TEC and found it increased realized matches, raised exposure per active template, reduced the share of templates stuck at low exposure, and improved both impression-level favoriting and downstream matching outcomes. In other words, the effect held up outside the simulator, on live traffic, with real employers and real workers responding to a redesigned ranking.
This Is Not a Timee-Specific Bug
What makes this finding worth generalizing is that the underlying pathology, popularity as a false proxy for exposure fairness, is a documented, recurring failure mode in recommender systems generally, not something unique to spot-work apps. Researchers Himan Abdollahpouri, Masoud Mansoury, Robin Burke, and Bamshad Mobasher formalized this in their widely cited study of popularity bias, showing that a small set of already-popular items get systematically over-represented in recommendations while the majority of items receive negligible exposure, and that algorithms which amplify this bias more heavily also produce worse calibration to what users actually want. Their finding, using two real-world datasets outside the labor-market context entirely, is the general form of what Timee's researchers found in the specific, high-stakes setting of gig work: optimize for predicted engagement and you will over-serve the crowd favorites at the expense of everything else, including the things people would prefer if you showed them.
Uber Eats hit the identical wall years earlier, in a three-sided marketplace of eaters, restaurants, and delivery partners rather than a two-sided labor market. Its engineering team documented that when they ranked restaurants purely by predicted eater conversion, new and lower-volume restaurants got starved of orders regardless of quality, because a ranking model trained on historical clicks has no impressions to work with for anything unproven. They found a direct, measurable trade-off between conversion rate and what they called marketplace fairness, exposing less-popular restaurants to more eaters, and built two separate fixes: a multi-armed bandit with an upper-confidence-bound score that temporarily boosts under-exposed restaurants and decays as they accumulate impressions, plus a multi-objective optimization layer using quadratic programming to cap how much conversion the platform is willing to sacrifice for balanced exposure. Two companies, two industries, a decade-and-a-continent apart, arrived at the same diagnosis: pure popularity ranking degrades the health of the supply side, and the fix requires an explicit, engineered counterweight rather than hoping the algorithm self-corrects.
The General Lesson for Any Marketplace With Expiring Inventory
Spot work is an extreme case of a much broader category: marketplaces where inventory expires and refills unpredictably, so that yesterday's popularity is a weak signal for today's availability. Event tickets, short-term rental listings, appointment slots, and same-day service bookings all share the same structure as Timee's job templates. A concert venue that was popular last month may have no seats left tonight; a rental market that trends toward certain buildings may have zero current vacancies in the buildings everyone already favorited, while a less-clicked building down the street sits half-empty. Any recommender trained to maximize predicted engagement on this kind of inventory will, by construction, keep bidding up exposure on the same high-turnover, low-capacity items while the long tail of genuinely available inventory goes unseen. The stakes scale with how time-sensitive and how economically consequential the match is, and in a labor market losing an estimated 2.6% of GDP to hiring shortfalls and staff turnover, according to Nikkei Asia's reporting on Japan's non-manufacturing sectors, misallocated exposure is not a UX nitpick. It is measurable economic waste.
Popularity tells you what people already wanted. It says nothing about what is currently open.
The operational takeaway is a single audit question, not a redesign mandate: does your ranking signal include a live measure of current capacity, or does it only measure historical engagement? If the answer is the latter, you already have Timee's original bug baked into your platform, whether or not your inventory looks anything like a part-time shift. The Timee researchers didn't need to abandon personalization or favoriting to fix it; they needed to add one more input, unfilled capacity, and gate the popularity signal against it. That is a targeted, measurable change, and on the numbers here, it is worth roughly twelve points of match rate.
Sources
- Sekiya, Otani, Komatsu, Fujii, Ozeki, and Noda, "Designing Recommendation Exposure and Favorite Lists: A Field Experiment in a Spot-Work Platform," arXiv preprint, 2026 arxiv.org
- "Exclusive-Japan spot work startup Timee targets July listing, sources say," Reuters (via Investing.com) investing.com
- "Shares of job app firm Timee jump in Japan stock trading debut," The Japan Times (Bloomberg) japantimes.co.jp
- "Japan suffers opportunity loss equivalent to 2.6% of GDP from labor crunch," Nikkei Asia asia.nikkei.com
- Himan Abdollahpouri, Masoud Mansoury, Robin Burke, and Bamshad Mobasher, "The Impact of Popularity Bias on Fairness and Calibration in Recommendation," arXiv arxiv.org
- "Food Discovery with Uber Eats: Recommending for the Marketplace," Uber Engineering Blog uber.com