Gaming
Screen capture from live play with every keystroke, mouse delta and camera pose aligned to the frame it happened on.
World models and agents that act through a UI: the frame, the input that produced the next frame, and where the camera was when it did.
- 8
- Delivery files
- 12 min
- Playable here
These are web previews, not the data. Each clip is downscaled to play in a browser. The delivery files are full resolution and carry every stream the rig recorded.
How this is captured
- Video
- 1920 x 1080 at 60 fps
- Titles
- 8 games, one session each
- Camera matrices
- Per-frame camera pose on 5 of 8 sessions
- On every frame
- Screen capture · Per-frame keystrokes · Semantic actions · Mouse delta
- Telemetry
- 27 columns per frame
- On some sessions
- Camera-to-world 4x4 · Pinhole intrinsics · Shared clock · Cross-agent visibility · Second viewport · Rerun recording
- Ships as
- .mp4 · .csv · session.json
- Also on some sessions
- key_bindings.json · key_binding.json · .rrd
The video is the easy half
A world model needs the action that produced the next frame. Below is the real input record from one Watch Dogs 2 session: 6,967 frames at 60 fps, every key and mouse delta held on each one.
Input state per frame
scrub the lanes
- MoveForward
- Sprint
- MoveLeft
- MoveRight
- Aim
- Fire
- Jump
- Cover
- Duration
- 116.1 s
- Frames
- 6,967
- Capture
- 1920 x 1080
- Telemetry columns
- 27
Two players, one clock
The clip is agent 12's viewport of a 2-player Grand Theft Auto V session. Both machines wrote to a shared clock, so the separation traced below is measured between them on the same frames, not estimated from this one view.
Distance between agents across 163.9s
- Agents
- 2
- Clock offset
- 933.5 ms
- In view
- 37%
What a game session carries
The footage is the smaller half. Every frame carries the keys held, the semantic action they map to, the mouse delta, and — where the title exposes it — the camera's position and orientation in the game world.
Grand Theft Auto V
RUB axes · 3 keys · 3 actions
Monster Hunter World
RFU axes · 1 keys · 4 actions
Borderlands 2
RFU axes · 3 keys · 3 actions
BeamNG.drive
RFU axes · 10 keys · 10 actions
Dragon's Dogma 2
RFU axes · 8 keys · 8 actions
Ground-plane projection of the camera position across the shipped window, normalised per title — five game worlds in five unit systems, so a shared extent would draw four of them as a dot. Filled dot starts, hollow dot ends. The axis convention is printed because it is load-bearing: right-forward-up puts the floor on x and y, right-up-back puts it on x and z, and a loader that assumes one plots the other title's route against its own height.
Just Cause 3, SnowRunner, The Outer Worlds record no camera pose. The rows are shipped with p, yaw and pitch null rather than interpolated, and the keys and actions are there in full.
What the player did · 8.0 s of one session
Every row of the shipped telemetry carries the keys held and the semantic action they map to, roughly every 20 ms. Consecutive rows with the same action are one span here — 480 marks would say "there was input"; these say what was done.
Grand Theft Auto V
10 spans · 3 distinct actions
move_forward → move_forward|move_left → move_forward → move_forward|move_left → …
Monster Hunter World
9 spans · 4 distinct actions
Forward|Sprint → Forward|Sprint|SecondaryAttack → Forward|Sprint → Forward|Sprint|PrimaryAttack → …
Just Cause 3
17 spans · 5 distinct actions
MoveForward → GrappleAim|MoveForward → MoveForward → GrappleAim|MoveForward → …
9,837 framesGrand Theft Auto V
two-player coop, agent 12 (host) · 181 MB
3,717 framesMonster Hunter World
single agent, Ancient Forest · 601 MB
4,595 framesJust Cause 3
single agent · 153 MB
7,253 framesSnowRunner
single agent · 161 MB
7,317 framesThe Outer Worlds
single agent · 220 MB
1,605 framesBorderlands 2
— · 96 MB video plus the recording
3,826 framesBeamNG.drive
— · 327 MB video plus the recording
3,832 framesDragon's Dogma 2
— · 377 MB video plus the recording
Browse the set
5 groups, 8 records.
Open one to play its clips and read the capture metadata. Every group ships the same way; what changes is the work in front of the camera.
Two ways into the full set
Previews on this page are open. Full resolution files, every camera, and the accompanying telemetry sit behind one of these.
Request the sample pack
Tell us the task, environment and modality you are training for. We send the matching sample files and the schema documentation.
Reply within one business day
Enter a passcode
If we have already spoken, your passcode opens the full sample set and the download links without another form.
Seven day session