Yuhang Wang

VLAlert-Bench: A Unified Benchmark for Driving-Alert Decision Making

University of South Florida · MOTIF-Lab
Dataset & benchmark · 2026
VLAlert pipeline

A unified per-tick benchmark for deciding when a driving system should stay SILENT, OBSERVE, or ALERT.

13,534
videos
192,892
1-s ticks
3
alert actions
6
source datasets

Overview

VLAlert-Bench integrates six driving-event datasets into a single per-tick prediction task with three actions — SILENT, OBSERVE, and ALERT. At each one-second tick a model sees the last 8 frames and predicts the appropriate alert action. The benchmark provides 13,534 videos across five splits, 192,892 one-second labeled ticks, per-frame action labels, and split manifests, and hosts the full ADAS-TO-Critic mp4 corpus directly. Source datasets include Nexar Collision, DoTA, DAD, DADA-2000, ADAS-TO-Critic, and the Kaggle Accident set. Released under CC-BY-4.0.

Demos

Per-tick alert decisions on real driving-event footage.

Demo 1
Alert decision on an approaching hazard.
Demo 3
Escalation from OBSERVE to ALERT.
Demo 9
A safety-critical collision sequence.

BibTeX

@misc{wang2026vlalert,
  title  = {VLAlert-Bench: A Unified Benchmark for Driving-Alert Decision Making},
  author = {Wang, Yuhang and Zhou, Hao},
  year   = {2026},
  howpublished = {Hugging Face Datasets},
  url    = {https://huggingface.co/datasets/HenryYHW/VLAlert}
}