All categories

Gaming

Screen capture from live play with every keystroke, mouse delta and camera pose aligned to the frame it happened on.

World models and agents that act through a UI: the frame, the input that produced the next frame, and where the camera was when it did.

8
Delivery files
12 min
Playable here
Two-player coop run · 1/8

These are web previews, not the data. Each clip is downscaled to play in a browser. The delivery files are full resolution and carry every stream the rig recorded.

How this is captured

Video
1920 x 1080 at 60 fps
Titles
8 games, one session each
Camera matrices
Per-frame camera pose on 5 of 8 sessions
On every frame
Screen capture · Per-frame keystrokes · Semantic actions · Mouse delta
Telemetry
27 columns per frame
On some sessions
Camera-to-world 4x4 · Pinhole intrinsics · Shared clock · Cross-agent visibility · Second viewport · Rerun recording
Ships as
.mp4 · .csv · session.json
Also on some sessions
key_bindings.json · key_binding.json · .rrd

The video is the easy half

A world model needs the action that produced the next frame. Below is the real input record from one Watch Dogs 2 session: 6,967 frames at 60 fps, every key and mouse delta held on each one.

Input state per frame

scrub the lanes

  • MoveForward
  • Sprint
  • MoveLeft
  • MoveRight
  • Aim
  • Fire
  • Jump
  • Cover
MouseDelta
Duration
116.1 s
Frames
6,967
Capture
1920 x 1080
Telemetry columns
27
Agent 12 / host

Two players, one clock

The clip is agent 12's viewport of a 2-player Grand Theft Auto V session. Both machines wrote to a shared clock, so the separation traced below is measured between them on the same frames, not estimated from this one view.

Distance between agents across 163.9s

Agents
2
Clock offset
933.5 ms
In view
37%

What a game session carries

The footage is the smaller half. Every frame carries the keys held, the semantic action they map to, the mouse delta, and — where the title exposes it — the camera's position and orientation in the game world.

Grand Theft Auto V

RUB axes · 3 keys · 3 actions

Monster Hunter World

RFU axes · 1 keys · 4 actions

Borderlands 2

RFU axes · 3 keys · 3 actions

BeamNG.drive

RFU axes · 10 keys · 10 actions

Dragon's Dogma 2

RFU axes · 8 keys · 8 actions

Ground-plane projection of the camera position across the shipped window, normalised per title — five game worlds in five unit systems, so a shared extent would draw four of them as a dot. Filled dot starts, hollow dot ends. The axis convention is printed because it is load-bearing: right-forward-up puts the floor on x and y, right-up-back puts it on x and z, and a loader that assumes one plots the other title's route against its own height.

Just Cause 3, SnowRunner, The Outer Worlds record no camera pose. The rows are shipped with p, yaw and pitch null rather than interpolated, and the keys and actions are there in full.

What the player did · 8.0 s of one session

Every row of the shipped telemetry carries the keys held and the semantic action they map to, roughly every 20 ms. Consecutive rows with the same action are one span here — 480 marks would say "there was input"; these say what was done.

Grand Theft Auto V

10 spans · 3 distinct actions

move_forward → move_forward|move_left → move_forward → move_forward|move_left → …

Monster Hunter World

9 spans · 4 distinct actions

Forward|Sprint → Forward|Sprint|SecondaryAttack → Forward|Sprint → Forward|Sprint|PrimaryAttack → …

Just Cause 3

17 spans · 5 distinct actions

MoveForward → GrappleAim|MoveForward → MoveForward → GrappleAim|MoveForward → …

9,837 frames

Grand Theft Auto V

two-player coop, agent 12 (host) · 181 MB

3,717 frames

Monster Hunter World

single agent, Ancient Forest · 601 MB

4,595 frames

Just Cause 3

single agent · 153 MB

7,253 frames

SnowRunner

single agent · 161 MB

7,317 frames

The Outer Worlds

single agent · 220 MB

1,605 frames

Borderlands 2

· 96 MB video plus the recording

3,826 frames

BeamNG.drive

· 327 MB video plus the recording

3,832 frames

Dragon's Dogma 2

· 377 MB video plus the recording

Browse the set

5 groups, 8 records.

Open one to play its clips and read the capture metadata. Every group ships the same way; what changes is the work in front of the camera.

Two ways into the full set

Previews on this page are open. Full resolution files, every camera, and the accompanying telemetry sit behind one of these.

Request the sample pack

Tell us the task, environment and modality you are training for. We send the matching sample files and the schema documentation.

Reply within one business day

Request access

Enter a passcode

If we have already spoken, your passcode opens the full sample set and the download links without another form.

Seven day session

Enter passcode