LeBoard

Methodology

Every Parquet row and every camera frame is decoded. Nothing is sampled. Each result is tied to the commit SHA that was scanned.

Hours

Measures
How much recorded robot time the dataset contains.
Formula
Σ episode timesteps ÷ fps ÷ 3600
Source
Episode lengths from the pinned episode metadata and fps from meta/info.json. Each timestep counts once, no matter how many cameras it has.
Not measured
How much of that time is useful for training.

Camera absent and camera dropout

Measures
Camera frames that are almost entirely one flat black or white value, split by whether the camera recorded anything in that episode.
Absent
An episode camera that is near-blank on every frame. The dataset has a slot for a camera that wasn’t recorded in that episode, usually filled with black. This is reported, not ranked. A camera that was dead for a whole episode looks the same and also counts as absent.
Dropout
A near-blank frame in an episode where the same camera has at least one live frame.
Formula
dropout frames ÷ (decoded camera frames − absent frames); absent share is absent frames ÷ decoded camera frames
Frame rule
≥ 99% clipped pixels, edge variance ≤ 1, mean brightness ≤ 2 or ≥ 253
Coverage
When video offsets can’t be placed in episodes, absence can’t be proven, so those near-blank frames count as dropouts. Scans before the split count every near-blank frame as a dropout.
Not measured
General image quality. Near-blank frames are excluded from blur, darkness and freeze.

Blur

Measures
Frames inside an interval where the camera is much blurrier than its own typical frame.
Formula
frames in blur intervals ÷ decoded camera frames
Rule
edge variance below 25% of this camera's nonblank median, for ≥ 0.5 s
Not measured
Whether the blur hurts training. Motion blur during fast moves is counted too.

Darkness

Measures
Frames inside an interval where most of the image is near-black with little contrast.
Formula
frames in darkness intervals ÷ decoded camera frames
Rule
at least 60% of pixels at luminance 16 or below, with contrast at most 40, for ≥ 0.5 s
Not measured
Scenes that are dark on purpose.

Freeze

Measures
Frames inside an interval where the video repeats the exact same image.
Formula
frames in freeze intervals ÷ decoded camera frames
Rule
consecutive pixel-exact repeated RGB frames, for ≥ 1 s
Not measured
A robot that is still while the camera keeps working. Small sensor noise breaks a freeze.

Duration mismatch

Measures
Episode-camera pairs whose metadata video interval disagrees with the Parquet frame count.
Formula
mismatched pairs ÷ assessed pairs
Rule
|interval − frames ÷ fps| > 1.01 nominal frames
Coverage
After one offset error, later pairs in the same file can’t be placed. They count as unassessed, not as matches.
Reading the cell
When some pairs are unassessed, the cell is hatched, the number gets a *, and assessed / total is shown under it. 0* means no mismatch among the pairs that could be checked, not a clean dataset. –* means the check ran but no pair could be assessed. A plain – means the check did not run.
Not measured
Whether visible motion lines up with recorded actions.

Flagged episodes

Measures
Distinct episodes flagged by any episode check: static actions, idle, non-finite values, episode boundaries or timestamp gaps. An episode with several reasons counts once.
Formula
flagged episodes ÷ assessed episodes; the dataset page also shows the hours inside flagged episodes
Coverage
Episodes shorter than 1 s or 10 steps, or whose values aren’t all finite, can’t be judged on motion and count as unassessed, so the cell is hatched like a partly assessed duration check. A check that can’t run on a dataset at all (idle, when it records no moving state) is left out of the count.
Not measured
Whether the remaining episodes are good for training. Video problems and absent cameras are not part of this count.

Static actions

Measures
Episodes whose recorded actions never change.
Rule
Every action dimension holds exactly the same value for the whole episode. There is no tolerance, so sensor noise never triggers it.
Coverage
Episodes shorter than 1 s or 10 steps, whichever is longer (max(ceil(fps), 10)), are unassessed: a short episode is trivially constant.
Not measured
Actions that change but are wrong, or that are stuck for only part of an episode.

Idle episodes

Measures
Episodes where the robot barely moves.
Scale
For each observation.state dimension, the median across the dataset’s episodes of how far that dimension moves within an episode. Dimensions whose scale is 0 (padding, an unused arm, a parked base) are left out.
Rule
Every remaining state dimension moves less than 2% of its scale. The comparison is per dimension, so radians, millimetres and gripper units mix safely.
Coverage
Episodes shorter than 1 s or 10 steps, or with non-finite values, are unassessed and left out of the scale. If no state dimension moves anywhere in the dataset (it records no state), the check is not run for that dataset.
Not measured
Idle stretches inside an episode that otherwise moves. The scale comes from the dataset itself, so if most episodes are idle the scale shrinks and idle episodes can go unflagged: this finds outliers within a dataset and doesn’t compare datasets.

Structure checks

Episode boundaries
Each Parquet row has the right episode_index, a contiguous global index and a frame_index counting from 0.
Finite action and state
Every action and observation.state value is a finite number.
Timestamp gaps
Timestamps increase, and no step is longer than 1.5 ÷ fps.
Video decoding
Every video file in the revision decodes.

Video/action sync

Status
Not measured. Matching durations and clean timestamps don’t show that visible motion lines up with the recorded actions.

Video files and times

LeRobot v3 stores many episodes back to back in one .mp4 per camera. Evidence Start and End are positions in that file (m:ss.s), which is where the video player seeks.

Missing a check? Request a metric →