Dataset Viewer
Auto-converted to Parquet Duplicate
Search is not available for this dataset
video
video
label
class label
0sample-0
1sample-1
2sample-10
3sample-100
4sample-101
5sample-102
6sample-103
7sample-104
8sample-105
9sample-106
10sample-107
11sample-108
12sample-109
13sample-11
14sample-110
15sample-111
16sample-112
17sample-113
18sample-114
19sample-115
20sample-116
21sample-117
22sample-118
23sample-119
24sample-12
25sample-120
26sample-121
27sample-122
28sample-123
29sample-124
30sample-125
31sample-126
32sample-127
33sample-128
34sample-129
35sample-13
36sample-130
37sample-131
38sample-132
39sample-133
40sample-134
41sample-135
42sample-136
43sample-137
44sample-138
45sample-139
46sample-14
47sample-140
48sample-141
49sample-142
50sample-143
51sample-144
52sample-145
53sample-146
54sample-147
55sample-148
56sample-149
57sample-15
58sample-150
59sample-151
60sample-152
61sample-153
62sample-154
63sample-155
64sample-156
65sample-157
66sample-158
67sample-159
68sample-16
69sample-160
70sample-161
71sample-162
72sample-163
73sample-164
74sample-165
75sample-166
76sample-167
77sample-168
78sample-169
79sample-17
80sample-170
81sample-171
82sample-172
83sample-173
84sample-174
85sample-175
86sample-176
87sample-177
88sample-178
89sample-179
90sample-18
91sample-180
92sample-181
93sample-182
94sample-183
95sample-184
96sample-185
97sample-186
98sample-187
99sample-188
End of preview. Expand in Data Studio

FastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams?

FastBench evaluates high-dynamic perception in streaming vision-language models. It contains 300 video clips and 306 English question-answer pairs, with six clips containing two questions. Models must observe video incrementally, capture brief events, and answer at the appropriate time.

This dataset accompanies the FastBench paper and its ProactiveFrame training-free adaptive frame-rate baseline. The release is an evaluation benchmark; no training or validation split is provided.

Benchmark coverage

Property Coverage
Temporal scopes Forward: 168 QA pairs; Instant: 73; Backward: 65
Domains Sports; Video Games; Performing Arts; Animals; Lifestyle & Recreation; Transportation; Science & Technology; Food & Cooking
Capabilities Action & Physical Interaction; Predictive & Causal Reasoning; Motion & Spatiotemporal Tracking; Entity & Visual Perception; Temporal & State Dynamics; Streaming & Online Detection
Evaluation Open-ended answers with an LLM judge and task-specific timing rules

The construction pipeline combines high-FPS QA generation, filtering out questions answerable at 2 FPS, trajectory-based verification using SAM3 and CoTracker3, and three rounds of human inspection. See the paper for the full protocol.

Files

.
├── README.md
├── proactive_perception_annotations.json
└── qa_video/
    ├── sample-0/original_cut.mp4
    ├── sample-1/original_cut1.mp4
    └── ...

The repository contains one annotation JSON file and 300 MP4 files, totaling approximately 8.3 GiB of video. Videos remain in their per-sample subdirectories. Local review sidecars (qa_annotation_checked.json) and internal metadata (annotation_status, remap_meta) are excluded. Questions, reference answers, temporal labels, and timestamps are preserved.

Download and use with the evaluator

Download this dataset repository into the dataset/ directory of the FastBench code repository. Replace YOUR_ACCOUNT/FastBench with this dataset's actual Hugging Face repository ID:

python -m pip install --upgrade huggingface_hub
cd /path/to/FastBench
hf download YOUR_ACCOUNT/FastBench --repo-type dataset --local-dir dataset

This produces dataset/proactive_perception_annotations.json and dataset/qa_video/. Annotation paths intentionally begin with dataset/qa_video/ and resolve relative to the code repository root, not the annotation file's directory. Keep these paths unchanged. When running the evaluator from the code repository root, set VIDEO_ROOT="$PWD".

Both config.yaml (baselines) and configs/proactive_frame.yaml (ProactiveFrame) in the code release use this annotation filename. See the code repository README for model deployment, evaluation, and judge configuration.

Annotation format

The JSON is a list of video records. Each record contains:

Field Meaning
video_path Relative video path, e.g. dataset/qa_video/sample-0/original_cut.mp4
verified_responses List of QA objects associated with this clip

Each QA object contains:

Field Meaning
user_query English question presented to the model
response Reference answer
time_type Temporal scope: forward, instant, or backward
timestamp_question Question time within the released clip
timestamp_proactive Target response time, when provided
timestamp_focus Annotated focus time, when provided; available for annotation-guided ablations
capability Evaluated perception or reasoning capability
domain Video domain label
*_seconds_in_original Optional timing metadata in seconds on the original source timeline

Clip timestamps are time strings (for example, 00:20). Original-source timing metadata uses a different timeline and should not replace the clip timestamps. The evaluator determines the precise question and response schedule from the temporal scope. A limit on video records can include more QA pairs than the limit itself.

For direct inspection without a dataset loader:

import json
from pathlib import Path

code_root = Path("/path/to/FastBench")
with (code_root / "dataset/proactive_perception_annotations.json").open(encoding="utf-8") as handle:
    records = json.load(handle)
video_file = code_root / records[0]["video_path"]
questions = records[0]["verified_responses"]

Intended use and limitations

FastBench is intended for research evaluation of streaming video perception. Its emphasis on brief, dynamic events and its size limit how broadly scores can be generalized. Report the model checkpoint, frame-sampling policy, context budget, judge configuration, and timing protocol when comparing results. Annotated focus times should not be used to trigger observations in the default ProactiveFrame setting.

License and attribution

The MIT license for the evaluation code does not grant rights to these videos. No dataset-wide license is declared in this card; annotation and media usage remains subject to applicable rights and source terms.

When using FastBench, cite FastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams? A formal citation can be added when the paper's public bibliographic details are available.

Downloads last month
56