What do Steam reviews actually praise and criticize and how much trust should you place in the answer? SteamLens is an evaluation-first review analysis system planned across four milestones, the first two delivered. It turns unstructured reviews into measured, aspect-level evidence, then builds the product around a labeler whose quality, cost, and failure modes are published.
A four-part research arc: a measurement instrument that had to earn trust first, then one question asked three times with the ground truth progressively removed — can learning rediscover optimal play as a table, as a network, and finally as a bet-sizing rule whose signal is buried fifty-deep in per-hand noise?
A learned A* heuristic that looked like a wash — until the pooled average was split and two opposite effects appeared. One regime-tag feature turns it into 17% less search at a 0.2% cost; moved off its training distribution, it fails in two opposite ways. Two axes, never collapsed into one score.
What does a Steam rating actually measure? 298k reviews, 50 games, 30 languages — four findings about when players review, whether they stay, how they write, and who they are, each forced to reproduce game-by-game before being believed. One 42-point "finding" died under that test; that's why the survivors can be trusted.