Simulation-grade gameplay data for world models, agents, and robotics. Rights-cleared at the source, captured frame-perfect from the running game, and normalised across a catalogue of 50+ licensed titles. Ground truth, not estimated after the fact.
Captured directly from the running game with PRISM. Cleared to train, disclosable under the EU AI Act.Most game data on the market is inferred: state read off the HUD with OCR, depth and normals generated after the fact by an estimator. That inference is also what generates the labels, under rules and definitions you never see. Every error poisons the dataset.
Ours is read directly from the running game. Labels are computed from that ground truth with explicit rules, through a shared game lexicon that normalises terminology across engines, transparently exposed and queryable.
Below, a real capture session with its signals playing in sync: scrub to any frame and read the exact state, input, and buffers behind it.
Real capture session: switch render buffers on the video, scrub to any frame and read the exact state behind it, all from the session's VTX file. Our homework, shown, not described.
Visual ground truth straight from the engine. Pixels and the render buffers behind them.
What the human did, synchronised with the state stream. The action half of any imitation-learning pair.
Everything the game is doing, frame by frame. The structured world behind the screen, not reconstructed from pixels.
Added labels generated after capture. Semantic, inferred, and consistent across every supported title.
+ added after captureEvery title is normalised to a single coordinate system, scale, and convention. Engine quirks removed. A grant spanning dozens of titles behaves like one dataset, not dozens of integrations, so you train and mix across the whole catalogue with zero per-title wrangling.
Skeletal bone data: hierarchical, per-joint rotations, where cross-engine normalisation usually breaks. Each joint's rotation is relative to its parent, so a single convention error compounds down the entire chain, and no two engines rig the same way. We resolve all of it to one standard. Captured at run-time from real player sessions, not replays.
From 2 August 2026 the EU AI Act obliges general-purpose AI providers to publish a summary of their training data; California's AB 2013 already requires disclosure. Gameplay cleared only at the player or platform layer is exactly what those summaries expose. Ours is licensed at the source, so you can name it and move on.
Captured only from titles we license directly, with the right to train. The studio copyright layer, the one in every frame, is cleared, not just the player's recording.
Every session's manifest carries its source, licence, and rights scope. Full chain of custody, per delivery, not a blanket warranty clause.
Uncleared audio and third-party assets identified and handled per title, up front, not discovered later at diligence.
Synthetic capture doesn't escape this: data generated from a game is still derived from that game's IP. The rights question follows the title, not the capture method. Ours arrives already answered.
A session is a folder you can open: standard mp4 for every visual signal, world state, inputs and events in VTX (our open, Apache-2.0 format), and one manifest describing it all. Install the Zenos Data CLI to search your grant, pull, and convert to Parquet or JSON. VTX spec on GitHub · CLI docs in the portal
Your schema changes, the data doesn't. Re-derive labels from exact state, no recapture, no re-licensing. The data outlives any single training run.
mp4 + VTX + one manifest per session. No proprietary reader between you and the data.
Every derivative traces back to the same cleared licence. The rights scope travels with the data.
Request access and the catalogue opens in the lab portal: titles, labels, and hours up front, with free samples to evaluate straight away. The data itself unlocks per order. Agree a volume of hours across the titles you want, each IP owner signs off on the use, and the data lands in your workspace. From there it's self-serve: search at full depth, pull, train. Every session carries its licence ID and rights scope in the manifest.
Register in the portal. Free samples to evaluate immediately.
Agree hours, titles, and signals with us.
Already licensed. Each IP owner confirms the use, built into the deal.
The data lands in your workspace.
Search, pull, convert. Self-serve from here.