Skip to content

Dataset & benchmark · driver motion · arXiv 2026

DriveMotionA Large-Scale Multi-Source Benchmark for Driver Motion Sequence Modeling and Forecasting

A multi-source, skeleton-only benchmark that forecasts the driver's whole body — 133 keypoints at 10 Hz — seconds ahead, anchored on real manoeuvres rather than on moments when nothing happens.

Three sources through one extraction trunk: a continuous fleet (time-aligned CAN, device head pose), a curated web corpus (front / side / back views, five visibility levels) and the AIDE clips (behaviour and emotion labels). Every sequence is released as a privacy-reduced skeleton video plus a standardized keypoint tensor — the behaviour is kept, the appearance is not.

arXiv 2609.08117 · 2026400 h · 9,010 sequences · 360 drivers (release)CC BY 4.0 · skeleton-derived

Yuhang Wang1·Chuheng Wei2·Jingxin Yang3·Xishun Liao4·Hao Zhou1

1University of South Florida · MOTIF-Lab   2Purdue University   3Stanford University   4University of Central Florida

Skeleton-only render of a driver in the fleet cabin camera
live · fleet cabin camera → 133 keypoints · 10 Hzskeleton only

Release (HF card, 23 Sep 2026): 400 h; 9,010 sequences = BATON 1,347 routes (320 h) + web 4,765 spans (78 h) + AIDE 2,898 clips (2.4 h). arXiv v1 (8 Sep 2026): 393 h; 4,367 forecasting sequences = 1,329 BATON routes (319 h) + 3,038 web spans (71.7 h), plus the 2,898 AIDE clips — the paper's sequence count excludes AIDE. Both: 360 drivers, 680,082 windows, 76,026 CAN-mined manoeuvre initiations.

01 · OVERVIEW

Forecast the body, not the face

What the benchmark is, in the paper's own words, beside what you need to know at a glance.

DriveMotion is a multi-source benchmark for continuous driver motion forecasting. It contains 393 hours of 133-keypoint motion sequences at 10 Hz from 360 drivers, integrating naturalistic driving data, curated public in-cabin videos, and the AIDE dataset into one skeleton representation. The primary protocol anchors evaluation on vehicle-dynamics transitions mined from CAN signals without providing CAN to the model at inference. Arm motion in pre-maneuver windows is 3.4× greater than in route-matched stable-driving controls. On these anchored windows, learned forecasters beat persistence by up to 15 %, while maneuver-enriched training improves forecast-derived Part-State F1 by 44 % over the zero-motion reference. Training on the full multi-source corpus further reduces forecasting error on held-out web drivers by 38 % compared with BATON-only training.

Condensed from the arXiv v1 abstract; sentences quoted verbatim carry a source chip.

02 · SKELETON PLAYER

One sequence format, three very different sources

Confidence sets the alpha of each joint; wrists and nose leave two-second trails; the ghost one second ahead is the ground truth the forecaster has to reach. Part LEDs show which body parts the extractor tracked; the CAN strip appears only where the fleet recorder had it.

Skeleton render of a fleet driver
Skeleton-only render (fleet cabin).

Fleet cabin fisheye, 40-s window chosen by brake and steering activity; six tracked parts; speed, steering and brake from the time-aligned CAN. The "this clip" meter divides the last two seconds of wrist speed by this clip's median — it is a live reading, not the paper's 3.4× statistic.

03 · EVIDENCE

What the numbers rest on

3.4× arm motion before a manoeuvre

The body moves before the car does.

Arm amplitude in pre-manoeuvre windows is 3.4× that of route-matched stable-driving controls (parked excluded), measured on ground truth over 76,026 CAN-mined manoeuvre initiations from 1,395 routes — which is why the benchmark anchors its primary protocol on these moments instead of random windows.

Motion energy around manoeuvre initiations
Motion energy around manoeuvre onsets (paper figure).
6.63 MPJPE@4 s · Transformer-L

Learned forecasters beat persistence — by up to 15 % — and enrichment lifts action-state F1 by 44 %.

Anchored board, zero-motion first: MPJPE at 4 s (bar, lower is better) and Part-State F1 at 2 s (second value, higher is better). The autoregressive token model has the highest F1 (.458) while the Transformer-L has the lowest MPJPE (6.63) — geometric and behavioural accuracy do not move together.

Anchored board in the paper.
−38 % error on held-out web drivers

Training on all three sources generalizes where fleet-only training does not.

Protocol B — robustness under observation shift: the full multi-source corpus reduces forecasting error on held-out web drivers by 38 % compared with BATON-only training. Viewpoint and visibility diversity are what the web corpus buys.

View-type counts in the release manifest.
9,010 sequences · 3 sources

One trunk, three sources, honest accounting.

Release manifest: BATON 1,347 routes (320 h, 301 drivers), web 4,765 spans (78 h, 23 creators), AIDE 2,898 clips (2.4 h). The paper's forecasting count (4,367) excludes the AIDE clips; the release counts them. Web curation funnel in the paper: 1,868 annotated videos → 873 driver videos → 79.9 h accepted → 71.7 h released.

Composition in the paper.
2,898 AIDE clips with labels

Semantics ride along in the same format.

AIDE's behaviour, emotion and scene labels are attached to the re-extracted sequences, so the player can show "Body Movement · Weariness · Smooth Traffic" next to the skeleton. Behaviour classes below; the format treats labels as metadata, not as supervision for the forecaster.

AIDE label counts.

What we do not claim. Skeletons are not identities and not emotions: the benchmark forecasts motion, and the AIDE labels are reported, not inferred. No source pixels of any person appear on this page or in the release.

04 · METHOD

From three kinds of video to one skeleton format

RTMW whole-body pose on every source, resampled to 10 Hz, with per-joint confidence, part validity, head pose and — for the fleet — time-aligned CAN; identity-disjoint splits; a dynamics-anchored protocol that never shows CAN to the model at inference.

DriveMotion pipeline from sources to skeleton sequences and the benchmark
Pipeline. Sources → extraction trunk → sequence format → protocols (release figure).
Protocol A: dynamics-anchored strata; Protocol B: observation shift
Protocols. A — pre-manoeuvre / post-manoeuvre / stable-control strata anchored on CAN events; B — held-out web drivers.
Composition statistics of the release
Composition. Sources, views and visibility levels (release figure).
Qualitative forecasts around manoeuvre onsets
Qualitative. Forecasts around manoeuvre onsets (skeletons only).
Diversity of DDPM samples
Stochastic forecasts. DDPM sample diversity.
05 · RESULTS

The anchored board

Protocol A · dynamics-anchored test windows · arXiv v1 · lower MPJPE is better, higher F1 is better

ModelMPJPE@4 s ↓Part-State F1@2 s ↑
Zero-motion (persistence)7.75.215
GRU6.83.275
siMLPe6.85.227
Transformer ED6.75.282
Transformer + context (enriched)6.95.309
Transformer-L6.63.287
Transformer-XL6.62.284
CVAE6.84.257
DDPM8.86.299
AR-LM (token model)8.16.458
Llama-3B8.24.457

Honest limits. Decoding under-expresses the predictive signal: the models with the strongest geometry are not the ones with the highest action-state F1. Per-frame forecasts were not retained for the release renders. The web corpus is skeleton-only by construction and its source terms are retained by the sources.

Sources & notes
    06 · DATA & ACCESS

    Get the sequences

    On request

    DriveMotion · full release

    9,010 sequences as keypoint tensors + skeleton videos, CAN and head pose where available, splits and protocol files. ~400 GB.

    as of 2026-10-02 · gated: manual · CC BY 4.0 (skeleton-derived)

    In the release

    Code & protocol

    Extraction trunk, sequence format spec (ODMS), baselines and the anchored protocol ship with the dataset repository.

    Not released

    Source video

    No cabin pixels: fleet cabin video stays with BATON's gated release, web footage stays with its creators, AIDE with its authors. Takedown requests are honoured.

    “Every sequence in DriveMotion is released as a privacy-reduced skeleton motion video plus a standardized keypoint tensor — the driver's behavior is preserved, appearance identity is not.”Hugging Face dataset card · fetched 2026-10-02
    07 · CITE

    Read the paper, cite the benchmark

    BibTeX
    @article{wang2026drivemotion,
      title   = {DriveMotion: A Large-Scale Multi-Source Benchmark for Driver Motion Sequence Modeling and Forecasting},
      author  = {Wang, Yuhang and Wei, Chuheng and Yang, Jingxin and Liao, Xishun and Zhou, Hao},
      journal = {arXiv preprint arXiv:2609.08117},
      year    = {2026},
      url     = {https://huggingface.co/datasets/HenryYHW/DriveMotion}
    }

    Related work from the lab