Dataset Viewer
The dataset viewer is not available for this subset.
Cannot get the split names for the config 'default' of the dataset.
Exception:    SplitsNotFoundError
Message:      The split names could not be parsed from the dataset config.
Traceback:    Traceback (most recent call last):
                File "/usr/local/lib/python3.12/site-packages/datasets/inspect.py", line 286, in get_dataset_config_info
                  for split_generator in builder._split_generators(
                                         ^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.12/site-packages/datasets/packaged_modules/json/json.py", line 97, in _split_generators
                  pa_table = next(iter(self._generate_tables(**splits[0].gen_kwargs, allow_full_read=False)))[1]
                             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.12/site-packages/datasets/packaged_modules/json/json.py", line 299, in _generate_tables
                  self._cast_table(pa_table, json_field_paths=json_field_paths),
                  ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.12/site-packages/datasets/packaged_modules/json/json.py", line 132, in _cast_table
                  pa_table = table_cast(pa_table, features.arrow_schema)
                             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.12/site-packages/datasets/table.py", line 2321, in table_cast
                  return cast_table_to_schema(table, schema)
                         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.12/site-packages/datasets/table.py", line 2255, in cast_table_to_schema
                  cast_array_to_feature(
                File "/usr/local/lib/python3.12/site-packages/datasets/table.py", line 1804, in wrapper
                  return pa.chunked_array([func(chunk, *args, **kwargs) for chunk in array.chunks])
                                           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.12/site-packages/datasets/table.py", line 2061, in cast_array_to_feature
                  casted_array_values = _c(array.values, feature.feature)
                                        ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.12/site-packages/datasets/table.py", line 1806, in wrapper
                  return func(array, *args, **kwargs)
                         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.12/site-packages/datasets/table.py", line 2014, in cast_array_to_feature
                  return pa.StructArray.from_arrays(arrays, names=list(feature), mask=array.is_null())
                         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "pyarrow/array.pxi", line 4295, in pyarrow.lib.StructArray.from_arrays
                File "pyarrow/array.pxi", line 1842, in pyarrow.lib.Array.validate
                File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status
              pyarrow.lib.ArrowInvalid: Struct child array #2 invalid: Invalid: Length spanned by list offsets (50501) larger than values array (length 50499)
              
              The above exception was the direct cause of the following exception:
              
              Traceback (most recent call last):
                File "/src/services/worker/src/worker/job_runners/config/split_names.py", line 66, in compute_split_names_from_streaming_response
                  for split in get_dataset_split_names(
                               ^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.12/site-packages/datasets/inspect.py", line 340, in get_dataset_split_names
                  info = get_dataset_config_info(
                         ^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.12/site-packages/datasets/inspect.py", line 291, in get_dataset_config_info
                  raise SplitsNotFoundError("The split names could not be parsed from the dataset config.") from err
              datasets.inspect.SplitsNotFoundError: The split names could not be parsed from the dataset config.

Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.

V-JEPA 2 ViT-G Embeddings — BEHAVIOR-1K 2025 Challenge Demos (62h)

Precomputed video embeddings for a 62-hour subsample of the BEHAVIOR-1K 2025 challenge demonstrations, extracted with the V-JEPA 2 ViT-g encoder. The goal is to make downstream experimentation faster and more reproducible by eliminating repeated video decoding and encoder forward passes — lowering the barrier for teams without access to large GPU clusters.

Field Value
Source dataset behavior-1k/2025-challenge-demos
Subset size ~62 hours
Rows ~1.1M
Encoder V-JEPA 2 ViT-g
Representation Precomputed video embeddings + proprioception + actions
Domain Embodied AI · Robotics · World Models
License MIT

Note: This repository contains precomputed feature representations only. It does not redistribute V-JEPA 2 source code or model weights.

Dataset Schema

Each row corresponds to one subsampled step/frame + actions/propios at 5 FPS from the original 30 FPS episodes.

Token columns — visual embeddings (one per camera view)

Column Type Shape Notes
tokens_head float16 ndarray (256, 1408) Head camera
tokens_left_wrist float16 ndarray (256, 1408) Left wrist camera
tokens_right_wrist float16 ndarray (256, 1408) Right wrist camera

256 spatial patches = (256px / 16px patch)² with tubelet size 2. 1408 = ViT-G embedding dim. float16 — bfloat16 is cast down at encoding time.

Proprioceptive columns — one entry per sampled frame

Column Type Shape Notes
actions float32 ndarray (138,) Flattened action chunk from frame f to f+1; fstp=6 at 30fps→5fps, 23 DoF → shape (6×23,)
states float32 ndarray (133,) Curated proprioceptive state at frame f
cam_rel_poses float32 ndarray (21,) 3 cameras × (position [3] + quaternion [4]) at frame f

Index columns — scalars

Column Type Description
frame_index int Source MP4 frame number
episode_idx int Episode index within the dataset
sample_idx int Manifest index for the episode
step_pos int 0-indexed position of this row within its episode
episode_len int Total rows in this episode

Dataset Pipeline

For details on dataset creation and processing, see the SHARP-Laps pipeline repository.

Licensing

This dataset is released under the MIT license.

Users should comply with the licenses of both the original BEHAVIOR-1K dataset and the V-JEPA 2 resources.

Citation

If you use this dataset, please cite BEHAVIOR-1K and V-JEPA 2 and V JEPA 2.1.

@misc{Quast2026,
  title={Short Horizon Planning with V-JEPA-2 AC on BEHAVIOR-1K},
  author={Quast, Julian},
  year={2026},
}

@article{li2024behavior,
  title={Behavior-1k: A human-centered, embodied ai benchmark with 1,000 everyday activities and realistic simulation},
  author={Li, Chengshu and Zhang, Ruohan and Wong, Josiah and ...},
  journal={arXiv preprint arXiv:2403.09227},
  year={2024}
}

@article{assran2025vjepa2,
  title={V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning},
  author={Assran, Mahmoud and Bardes, Adrien and Fan, David and ...},
  journal={arXiv preprint arXiv:2506.09985},
  year={2025}
}

@article{murlabadia2026vjepa2_1,
  title={V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning},
  author={Mur-Labadia, Lorenzo and Muckley, Matthew and Bar, Amir and ...},
  journal={arXiv preprint arXiv:2603.14482},
  year={2026}
}
Downloads last month
1,214

Papers for quastAI/behavior-1k-2025-challenge-vjepa2-vitg-demo-embeddings