Datasets:
The dataset viewer is not available for this subset.
Exception: SplitsNotFoundError
Message: The split names could not be parsed from the dataset config.
Traceback: Traceback (most recent call last):
File "/usr/local/lib/python3.12/site-packages/datasets/inspect.py", line 286, in get_dataset_config_info
for split_generator in builder._split_generators(
^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/site-packages/datasets/packaged_modules/json/json.py", line 97, in _split_generators
pa_table = next(iter(self._generate_tables(**splits[0].gen_kwargs, allow_full_read=False)))[1]
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/site-packages/datasets/packaged_modules/json/json.py", line 299, in _generate_tables
self._cast_table(pa_table, json_field_paths=json_field_paths),
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/site-packages/datasets/packaged_modules/json/json.py", line 132, in _cast_table
pa_table = table_cast(pa_table, features.arrow_schema)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/site-packages/datasets/table.py", line 2321, in table_cast
return cast_table_to_schema(table, schema)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/site-packages/datasets/table.py", line 2255, in cast_table_to_schema
cast_array_to_feature(
File "/usr/local/lib/python3.12/site-packages/datasets/table.py", line 1804, in wrapper
return pa.chunked_array([func(chunk, *args, **kwargs) for chunk in array.chunks])
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/site-packages/datasets/table.py", line 2061, in cast_array_to_feature
casted_array_values = _c(array.values, feature.feature)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/site-packages/datasets/table.py", line 1806, in wrapper
return func(array, *args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/site-packages/datasets/table.py", line 2014, in cast_array_to_feature
return pa.StructArray.from_arrays(arrays, names=list(feature), mask=array.is_null())
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "pyarrow/array.pxi", line 4295, in pyarrow.lib.StructArray.from_arrays
File "pyarrow/array.pxi", line 1842, in pyarrow.lib.Array.validate
File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status
pyarrow.lib.ArrowInvalid: Struct child array #2 invalid: Invalid: Length spanned by list offsets (50501) larger than values array (length 50499)
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
File "/src/services/worker/src/worker/job_runners/config/split_names.py", line 66, in compute_split_names_from_streaming_response
for split in get_dataset_split_names(
^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/site-packages/datasets/inspect.py", line 340, in get_dataset_split_names
info = get_dataset_config_info(
^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.12/site-packages/datasets/inspect.py", line 291, in get_dataset_config_info
raise SplitsNotFoundError("The split names could not be parsed from the dataset config.") from err
datasets.inspect.SplitsNotFoundError: The split names could not be parsed from the dataset config.Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.
V-JEPA 2 ViT-G Embeddings — BEHAVIOR-1K 2025 Challenge Demos (62h)
Precomputed video embeddings for a 62-hour subsample of the BEHAVIOR-1K 2025 challenge demonstrations, extracted with the V-JEPA 2 ViT-g encoder. The goal is to make downstream experimentation faster and more reproducible by eliminating repeated video decoding and encoder forward passes — lowering the barrier for teams without access to large GPU clusters.
| Field | Value |
|---|---|
| Source dataset | behavior-1k/2025-challenge-demos |
| Subset size | ~62 hours |
| Rows | ~1.1M |
| Encoder | V-JEPA 2 ViT-g |
| Representation | Precomputed video embeddings + proprioception + actions |
| Domain | Embodied AI · Robotics · World Models |
| License | MIT |
Note: This repository contains precomputed feature representations only. It does not redistribute V-JEPA 2 source code or model weights.
Dataset Schema
Each row corresponds to one subsampled step/frame + actions/propios at 5 FPS from the original 30 FPS episodes.
Token columns — visual embeddings (one per camera view)
| Column | Type | Shape | Notes |
|---|---|---|---|
tokens_head |
float16 ndarray |
(256, 1408) |
Head camera |
tokens_left_wrist |
float16 ndarray |
(256, 1408) |
Left wrist camera |
tokens_right_wrist |
float16 ndarray |
(256, 1408) |
Right wrist camera |
256 spatial patches = (256px / 16px patch)² with tubelet size 2. 1408 = ViT-G embedding dim.
float16 — bfloat16 is cast down at encoding time.
Proprioceptive columns — one entry per sampled frame
| Column | Type | Shape | Notes |
|---|---|---|---|
actions |
float32 ndarray |
(138,) |
Flattened action chunk from frame f to f+1; fstp=6 at 30fps→5fps, 23 DoF → shape (6×23,) |
states |
float32 ndarray |
(133,) |
Curated proprioceptive state at frame f |
cam_rel_poses |
float32 ndarray |
(21,) |
3 cameras × (position [3] + quaternion [4]) at frame f |
Index columns — scalars
| Column | Type | Description |
|---|---|---|
frame_index |
int |
Source MP4 frame number |
episode_idx |
int |
Episode index within the dataset |
sample_idx |
int |
Manifest index for the episode |
step_pos |
int |
0-indexed position of this row within its episode |
episode_len |
int |
Total rows in this episode |
Dataset Pipeline
For details on dataset creation and processing, see the SHARP-Laps pipeline repository.
Licensing
This dataset is released under the MIT license.
- Source demonstrations:
behavior-1k/2025-challenge-demos— MIT - V-JEPA 2 (Meta / FAIR): facebookresearch/vjepa2 — Apache-2.0 (not redistributed)
Users should comply with the licenses of both the original BEHAVIOR-1K dataset and the V-JEPA 2 resources.
Citation
If you use this dataset, please cite BEHAVIOR-1K and V-JEPA 2 and V JEPA 2.1.
@misc{Quast2026,
title={Short Horizon Planning with V-JEPA-2 AC on BEHAVIOR-1K},
author={Quast, Julian},
year={2026},
}
@article{li2024behavior,
title={Behavior-1k: A human-centered, embodied ai benchmark with 1,000 everyday activities and realistic simulation},
author={Li, Chengshu and Zhang, Ruohan and Wong, Josiah and ...},
journal={arXiv preprint arXiv:2403.09227},
year={2024}
}
@article{assran2025vjepa2,
title={V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning},
author={Assran, Mahmoud and Bardes, Adrien and Fan, David and ...},
journal={arXiv preprint arXiv:2506.09985},
year={2025}
}
@article{murlabadia2026vjepa2_1,
title={V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning},
author={Mur-Labadia, Lorenzo and Muckley, Matthew and Bar, Amir and ...},
journal={arXiv preprint arXiv:2603.14482},
year={2026}
}
- Downloads last month
- 1,214