Why this benchmark?
Holistic gesture quality is not one thing:
naturalness, diversity, synchrony, physical plausibility, and semantic appropriateness can move in different directions.
Recent holistic systems synthesize coordinated face, hand, upper-body, lower-body, and global motion, often conditioned on both speech acoustics and linguistic content.
Yet evaluation has not kept pace. Different works use different datasets, preprocessing, rendering pipelines, metrics, and subjective protocols, making reported gains hard to compare directly.
Even when models are evaluated under the same protocol, commonly used objective metrics may still fail to reflect human perception, particularly for semantic gestures that are sparse,
context-dependent, and can be expressed through many valid physical forms.
01
Heterogeneous protocols
Prior systems are often evaluated with incompatible training regimes, data subsets, sequence boundaries, renderers, and metric implementations.
Apparent model improvements may therefore reflect experimental choices rather than algorithmic progress.
02
Metrics vs. people
Objective measures are reproducible and scalable, but strong correspondence with human judgments is not guaranteed.
Our benchmark explicitly asks which metrics align with which perceptual properties instead of assuming that all scores are equally meaningful.
03
The semantic gap
A gesture can be smooth, plausible, and synchronized while still being semantically empty or mismatched.
Reference-pose metrics also penalize alternative gestures that may communicate the same meaning through different motion.