← All posts / Research

5,000 Hours of Elden Ring and Valorant: Tencent's GameHorizon Wants to Be the Yardstick for Game-Playing AI

Tencent ARC Lab's GameHorizon Suite benchmarked 47 models across 21 AAA titles and one million-plus evaluations — GPT-6 Astra leads at 80.2%, but every model struggles most with short-horizon action control.

5,000 Hours of Elden Ring and Valorant: Tencent's GameHorizon Wants to Be the Yardstick for Game-Playing AI

How well can an AI actually play a video game — not summarize its wiki, not generate fan art of it, but read the screen, decompose a goal, and land the inputs? For years that question has been answered with bespoke setups: one game, one agent, one paper, no way to compare results. Tencent’s ARC Lab is trying to end that fragmentation with GameHorizon Suite, a data-and-evaluation package published September 21 that spans 21 AAA titles, 5,000 hours of expert human gameplay, and a reproducible benchmark that has already been run against 47 models with more than one million model invocations.

What GameHorizon actually is

The suite has three components, and the design choices in each are what make it more than another gameplay dataset.

GameHorizon-Annotator is a scalable, automated pipeline that turns raw gameplay trajectories into a three-level pyramid of instructions: L1 short-horizon operations, L2 medium-horizon goals, and L3 long-horizon strategies. Videos and action streams first go through action-aware segmentation, which uses key actions to determine clip boundaries. A vision-language model annotates each clip with an L1 operation; a bottom-up merging procedure then progressively fuses lower-level clips into medium- and long-horizon segments based on action continuity and semantic coherence, from which the VLM derives L2 goals and L3 strategies. The pyramid structure — primitive actions, operations, goals, strategies — is what allows evaluation at multiple temporal horizons, the suite’s namesake feature.

GameHorizon-Data is the resulting corpus: the first large-scale AAA gameplay dataset with temporally aligned video recordings, player actions, and multi-horizon instructions, collected from 100 experienced human players. The 21 titles span five genres — open-world, action role-playing, competitive shooter, sandbox survival, and creature-collecting adventure. The roster reads like a Steam top-sellers list: Black Myth: Wukong, Elden Ring, Cyberpunk 2077, Red Dead Redemption 2, GTA V, The Witcher 3, Assassin’s Creed, Apex Legends, PUBG, Palworld, Escape from Tarkov, Genshin Impact, Valorant, Minecraft, Wuthering Waves, and more. Annotation density is substantial: Valorant alone contributes roughly 617 recorded hours yielding over 705,000 L1 instructions, 23,000 L2 goals, and 5,700 L3 strategies, with average instruction spans of 2.7 seconds at L1 rising to 332.5 seconds at L3.

GameHorizon-Bench is the evaluation layer, with two tracks. The offline track uses roughly 5,000 standardized multiple-choice questions across three primary tasks plus ten diagnostic variants: single-horizon action (T1, pick the correct action sequence for an L1 instruction), multi-horizon decomposition (T2, decompose an L2 goal into the right sequence of L1 operations), and cross-horizon consistency (T3, overall coherence across frames, L1–L3 instructions, and actions). The online track — 20 tasks decomposed into 62 verifiable subtasks — tests whether those offline scores translate into actual gameplay, splitting long-horizon challenges into causal tasks that must be completed in a prescribed order and thematic tasks whose subtasks share a theme but can be done in any order. Because each subtask is verifiable, the online track doubles as a failure-localization tool: when an agent stalls, you can see exactly which step broke.

What 47 models and a million invocations revealed

The headline leaderboard is for the offline primary tasks, and it reshuffles some expectations. GPT-6 Astra leads at 80.2% overall, with Gemini 3.8 Flash second at 77.3% and Gemini 3.7 Flash third at 76.7%. GPT-5.6 Sol sits fifth at 74.8%, just ahead of Kimi-K3 at 74.5%. Doubao-Seed-2.1-Turbo and -Pro both crack the top eleven. Claude Fable 5 lands mid-pack in Tier 2 at 71.2%, one spot ahead of GPT-5.6 Terra, while Claude Opus 4.8 sits in Tier 3 at 64.8% — a reminder that general-purpose reasoning strength doesn’t automatically confer fine-grained visuomotor competence.

The task-level averages are the more interesting finding. Averaged across all 47 models, accuracy is 57.3% on T1 (short-horizon action), 65.1% on T2 (decomposition), and 71.6% on T3 (consistency). In other words, models are comparatively good at judging whether a plan hangs together, mediocre at breaking goals into ordered operations, and weakest at the very thing human players do without thinking: choosing the correct moment-to-moment action sequence. The gap between frontier and small models is wide but not clean — GPT-4o posts a Tier 4 overall of 54.4%, and specialized GUI agents like UI-TARS-1.5-7B manage 53.0%, while the compact InternVL-U-4B trails at 44.6%.

Why it matters

Gameplay is one of the few domains where visual understanding, instruction decomposition, goal planning, and precise action control must all work simultaneously, under partial observability, with consequences that compound over time. That makes it a plausible proxy for the embodied and agentic tasks industry actually cares about — and a harder one than most static benchmarks, because “understanding the scene” and “acting correctly in it” are measured separately.

The existing landscape pushed the team toward this design: prior datasets either covered a narrow range of games, lacked language instructions, or relied on high-variance online rollouts that made results hard to reproduce. GameHorizon’s answer — reproducible offline MCQs plus verifiable stepwise online tasks — trades some ecological realism for comparability, which is exactly what a “standardized yardstick” (the authors’ phrase) requires. The team says dataset, annotator, and benchmark will all be released publicly.

The competitive stakes are not subtle. A Tencent lab publishing a benchmark where its own Doubao models place respectably, Google’s Flash line ranks second and third, and Western frontier models top the table is also a statement about who is measuring whom. But the dataset itself — five genres, 21 games, temporally aligned actions and instructions — is a genuine contribution for anyone training or evaluating game-playing agents, and the failure-localization design means the benchmark teaches you where models break, not just that they do.

The open question is whether offline MCQ accuracy will predict online gameplay success as cleanly as the field hopes. The paper’s own online track exists precisely to test that mapping, and the tier structure suggests the correlation will hold at the top but get noisy in the middle of the pack. Either way, with 47 models already scored and the release promised as open, GameHorizon is positioned to become the default answer to “how good is your agent at actually playing?”