IP Toolkit · Advisory

Training on games, without the legal risk.

Gameplay is the best training data on earth for spatial and embodied intelligence. It is also the most layered rights object in media. Here is the map, before disclosure makes it everyone's business.

For AI labs · world models · agents · robotics. General information, not legal advice.
Disclosure · what's already law
EU AI Act Art. 53summary mandatory
enforceable from2 Aug 2026
fines up to€15M / 3% turnover
California AB 2013in force since Jan 2026
who reads itevery rights-holder
The clock
53

From 2 August 2026, your training data is public information.

EU AI Act Article 53 obliges every general-purpose AI provider serving the EU to publish a training-content summary on the Commission's template, and to maintain a policy respecting rights-holders' text-and-data-mining reservations, wherever the training happened. From 2 August 2026 the AI Office can enforce, with fines up to EUR 15M or 3% of global turnover. California's AB 2013 has required disclosure since January 2026. Rights-holders will read what you publish. This guide is about being able to publish it without flinching.

The rights stack

One clip. Five layers of rights.

A single gameplay clip carries more rights layers than nearly any other training input. Whoever gave you the clip could, at most, give you the layers they held:

L5 · Engine + middleware

Structured state data, replays, and telemetry can implicate engine and middleware terms in addition to the studio's rights.

L4 · Platform ToS

Clip and streaming platforms take user uploads under terms claiming broad content rights. Those terms convey, at best, the player layer. Studios are now amending creator policies to state explicitly that AI-training rights were never the player's to pass on.

L3 · Player layer

The player's recording, inputs, voice, and personal data. Player consent covers this layer only: a player cannot grant rights to the game content itself.

L2 · Music + audio

Licensed soundtracks and voice work often carry third-party rights the studio itself cannot pass on. Uncleared audio is the most common carve-out in legitimate licences.

L1 · Studio copyright

The game's audio-visual output, characters, environments, UI, and assets belong to the studio or publisher. This layer exists in every frame of every clip, no matter who recorded it.

In every frame

The structural problem with clip-sourced corpora: platform terms and player opt-ins clear the thinnest layer of the stack. The studio copyright layer, the one in every pixel, is the one they do not touch.

The legal weather

What the recent cases actually established.

Nobody can tell you unlicensed training on games is unlawful everywhere; nobody serious can tell you it is safe either. The signals so far:

SignalWhat happenedWhat it means for games
Music's arc completedLabels sued Suno and Udio on output evidence (2024); both settled into licensing deals with major labels (2025).The pattern for media verticals: enforcement pressure converts grey training into licences. Games is the next vertical with organised rights-holders.
The market-harm splitOne US court recognised harm to a training-data licensing market (Ross); another called such a market hypothetical (Bartz); a third dismissed for lack of market-harm evidence while inviting better (Kadrey).The open question is whether a real, priced market exists. For game data, one now does: licensed gameplay catalogues sell at published rates for AI and robotics training. That evidence strengthens every future rights-holder claim.
Acquisition mattersThe largest AI copyright settlement to date ($1.5B) turned on how the training data was obtained, not on whether training is transformative.Provenance is not a detail. How your gameplay corpus was assembled may matter more than what you trained.
Outputs are exhibitsVideo models publicly reproduced recognisable game content: gameplay styles, title screens, characters.Distinctive game IP is exactly what memorisation probes find. Assume your model can be tested from the outside, because it can.
Clean provenance

The rights story you want behind every hour.

If you had to attach a rights story to every hour of gameplay in your corpus, the defensible version reads like this:

  1. 01

    Chain of title, both sides

    A licence from the studio or publisher covering AI training on the game's content, and consent covering the player layer, for the same footage.

  2. 02

    Carve-outs documented

    Uncleared audio and third-party assets identified and excluded or delivered in stripped versions, not discovered later.

  3. 03

    Named, disclosable sources

    A supplier and licence chain you can put in an Article 53 summary without creating a target on your own back.

  4. 04

    Terms that survive scale

    Sublicensing, retention, deletion, and model-survival terms agreed upfront, so a settled legal position does not unravel at your Series C diligence.

  5. 05

    No laundering steps

    No layer of the chain where the answer to "who granted the AI-training right?" is a platform ToS or an individual who never held it.

Due diligence

Eight questions for any gameplay data supplier.

01 · The licence

Do you hold a written licence from the studio or publisher of each title, expressly covering AI training?

02 · The player layer

Is the player layer cleared for the same footage, and how?

03 · Carve-outs

How are audio and third-party asset carve-outs identified and handled?

04 · Disclosability

Can we name you and your licence chain in a public training-content disclosure?

05 · Withdrawal

What happens to models trained on the data if a title is later withdrawn?

06 · Documentation

Do you provide per-delivery provenance documentation we can file, not just a warranty clause?

07 · Stability

Are the studio's revenue terms transparent, so the licence is stable rather than resented?

08 · The backstop

If a rights-holder challenges our use of your data, what stands behind your answer?

If a supplier hesitates on the first question, every other answer is decoration.

The clean path

Rights-cleared game data is not scarcer than grey data. It is better.

Engine-level capture beats screen scrapes, structured state and input data beats video alone, and made-to-order capture beats whatever happens to be on a clip platform. The licensed route costs money and removes the two things no lab can engineer away: litigation overhang and a disclosure you would rather not publish.

Where ZENOS fits. ZENOS operates a licensed gameplay data catalogue: studio and player cleared, carve-outs tracked per title, delivered with provenance documentation designed to slot into an Article 53 summary. Video, synchronised game state, and player inputs, from a growing catalogue of licensed titles, any volume made to order. If you are building world models, agents, or robotics policies on game data, we are the supply chain that lets you disclose it proudly.

General information, not legal advice. AI, copyright, and data law differ by jurisdiction and are developing quickly; take advice from your own counsel on any specific plan.

Build on real worlds.

Building frontier AI?Talk to us →
Have a catalogue?List your game →