Gameplay is the best training data on earth for spatial and embodied intelligence. It is also the most layered rights object in media. Here is the map, before disclosure makes it everyone's business.
For AI labs · world models · agents · robotics. General information, not legal advice.EU AI Act Article 53 obliges every general-purpose AI provider serving the EU to publish a training-content summary on the Commission's template, and to maintain a policy respecting rights-holders' text-and-data-mining reservations, wherever the training happened. From 2 August 2026 the AI Office can enforce, with fines up to EUR 15M or 3% of global turnover. California's AB 2013 has required disclosure since January 2026. Rights-holders will read what you publish. This guide is about being able to publish it without flinching.
A single gameplay clip carries more rights layers than nearly any other training input. Whoever gave you the clip could, at most, give you the layers they held:
Structured state data, replays, and telemetry can implicate engine and middleware terms in addition to the studio's rights.
Clip and streaming platforms take user uploads under terms claiming broad content rights. Those terms convey, at best, the player layer. Studios are now amending creator policies to state explicitly that AI-training rights were never the player's to pass on.
The player's recording, inputs, voice, and personal data. Player consent covers this layer only: a player cannot grant rights to the game content itself.
Licensed soundtracks and voice work often carry third-party rights the studio itself cannot pass on. Uncleared audio is the most common carve-out in legitimate licences.
The game's audio-visual output, characters, environments, UI, and assets belong to the studio or publisher. This layer exists in every frame of every clip, no matter who recorded it.
In every frameThe structural problem with clip-sourced corpora: platform terms and player opt-ins clear the thinnest layer of the stack. The studio copyright layer, the one in every pixel, is the one they do not touch.
Nobody can tell you unlicensed training on games is unlawful everywhere; nobody serious can tell you it is safe either. The signals so far:
| Signal | What happened | What it means for games |
|---|---|---|
| Music's arc completed | Labels sued Suno and Udio on output evidence (2024); both settled into licensing deals with major labels (2025). | The pattern for media verticals: enforcement pressure converts grey training into licences. Games is the next vertical with organised rights-holders. |
| The market-harm split | One US court recognised harm to a training-data licensing market (Ross); another called such a market hypothetical (Bartz); a third dismissed for lack of market-harm evidence while inviting better (Kadrey). | The open question is whether a real, priced market exists. For game data, one now does: licensed gameplay catalogues sell at published rates for AI and robotics training. That evidence strengthens every future rights-holder claim. |
| Acquisition matters | The largest AI copyright settlement to date ($1.5B) turned on how the training data was obtained, not on whether training is transformative. | Provenance is not a detail. How your gameplay corpus was assembled may matter more than what you trained. |
| Outputs are exhibits | Video models publicly reproduced recognisable game content: gameplay styles, title screens, characters. | Distinctive game IP is exactly what memorisation probes find. Assume your model can be tested from the outside, because it can. |
If you had to attach a rights story to every hour of gameplay in your corpus, the defensible version reads like this:
A licence from the studio or publisher covering AI training on the game's content, and consent covering the player layer, for the same footage.
Uncleared audio and third-party assets identified and excluded or delivered in stripped versions, not discovered later.
A supplier and licence chain you can put in an Article 53 summary without creating a target on your own back.
Sublicensing, retention, deletion, and model-survival terms agreed upfront, so a settled legal position does not unravel at your Series C diligence.
No layer of the chain where the answer to "who granted the AI-training right?" is a platform ToS or an individual who never held it.
Do you hold a written licence from the studio or publisher of each title, expressly covering AI training?
Is the player layer cleared for the same footage, and how?
How are audio and third-party asset carve-outs identified and handled?
Can we name you and your licence chain in a public training-content disclosure?
What happens to models trained on the data if a title is later withdrawn?
Do you provide per-delivery provenance documentation we can file, not just a warranty clause?
Are the studio's revenue terms transparent, so the licence is stable rather than resented?
If a rights-holder challenges our use of your data, what stands behind your answer?
If a supplier hesitates on the first question, every other answer is decoration.
Engine-level capture beats screen scrapes, structured state and input data beats video alone, and made-to-order capture beats whatever happens to be on a clip platform. The licensed route costs money and removes the two things no lab can engineer away: litigation overhang and a disclosure you would rather not publish.
Where ZENOS fits. ZENOS operates a licensed gameplay data catalogue: studio and player cleared, carve-outs tracked per title, delivered with provenance documentation designed to slot into an Article 53 summary. Video, synchronised game state, and player inputs, from a growing catalogue of licensed titles, any volume made to order. If you are building world models, agents, or robotics policies on game data, we are the supply chain that lets you disclose it proudly.