Holistic Co-Speech Gesture Evaluation

Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation

A perceptually grounded and semantics-aware benchmark for holistic co-speech gesture generation that combines standardized model comparison, human-centered metric validation, and fine-grained semantic evaluation.

Project webpage
4representative holistic gesture generation systems evaluated under one protocol
13objective metrics curated for perceptual correlation and composite analysis across distribution, geometry, motion quality, alignment, and semantics
101valid participants after quality control across muted and audio studies
265shared BEAT2 test sequences used for objective evaluation, except SGP
Why this benchmark?
Holistic gesture quality is not one thing: naturalness, diversity, synchrony, physical plausibility, and semantic appropriateness can move in different directions.

Recent holistic systems synthesize coordinated face, hand, upper-body, lower-body, and global motion, often conditioned on both speech acoustics and linguistic content. Yet evaluation has not kept pace. Different works use different datasets, preprocessing, rendering pipelines, metrics, and subjective protocols, making reported gains hard to compare directly. Even when models are evaluated under the same protocol, commonly used objective metrics may still fail to reflect human perception, particularly for semantic gestures that are sparse, context-dependent, and can be expressed through many valid physical forms.

01

Heterogeneous protocols

Prior systems are often evaluated with incompatible training regimes, data subsets, sequence boundaries, renderers, and metric implementations. Apparent model improvements may therefore reflect experimental choices rather than algorithmic progress.

02

Metrics vs. people

Objective measures are reproducible and scalable, but strong correspondence with human judgments is not guaranteed. Our benchmark explicitly asks which metrics align with which perceptual properties instead of assuming that all scores are equally meaningful.

03

The semantic gap

A gesture can be smooth, plausible, and synchronized while still being semantically empty or mismatched. Reference-pose metrics also penalize alternative gestures that may communicate the same meaning through different motion.

01 · Benchmark framework

Standardized training, inference, rendering, and evaluation.

The benchmark compares EMAGE, GestureLSM, SemTalk, and SemConFlow under a shared BEAT2 evaluation setting. The goal is not only to rank systems, but to make those rankings interpretable by controlling the sources of experimental variation.

Unified benchmark protocol

Common training regimeofficial implementations and optimization settings, trained with shared BEAT2 data
→
Common inferencesame test sequences, speech inputs, and sequence boundaries
→
Common evaluationSMPL-X conversion, post-processing, rendering, metrics, and perception study

Dataset & implementation

BEAT2

≈60 hours, 25 speakers, synchronized speech and holistic SMPL-X motion.

Split

Official 85% / 7.5% / 7.5% train / validation / test split across all 25 speakers.

Compute

All benchmark experiments are conducted on a single NVIDIA A100 GPU.

EMAGE

A unified holistic framework that jointly synthesizes facial expressions, upper- and lower-body motion, hands, and global translation. It uses a Masked Audio Gesture Transformer and four compositional VQ-VAEs for different body regions. Content Rhythm Attention adaptively combines speech rhythm and transcript semantics.

Masked modelingCompositional VQ-VAEsCRA

GestureLSM

A flow-matching framework for real-time holistic gesture generation. Region-specific residual vector-quantized tokens are coordinated through spatial and temporal attention, while a latent shortcut model reduces the number of sampling steps.

Flow matchingSpatial-temporal attentionLatent shortcut

SemTalk

Separates rhythm-related base motion from sparse semantic motion. A hierarchical coarse-to-fine module establishes rhythmic consistency, while semantic emphasis learning activates content-related motion at key frames using text, CLIP, emotion, and speech features.

Rhythm baseSparse semanticsFrame-level emphasis

SemConFlow

Learns region-specific RVQ-VAE motion priors and aligns the resulting holistic motion latent with text and audio. Contrastive flow matching uses mismatched audio-text contexts as negatives, encouraging semantically congruent trajectories and discouraging generic rhythm-dominated gestures.

RVQ-VAESemantic alignmentContrastive flow matching
02 · Objective evaluation

A multidimensional taxonomy of gesture quality.

The benchmark organizes objective evaluation by what each metric is intended to capture. This distinction is important because several metrics that look similar numerically measure fundamentally different properties.

Distributional similarity

  • FGD compares reference and generated distributions in a learned gesture feature space.
  • Density measures whether generated samples lie in well-supported regions of the real-data manifold.
  • Coverage measures how extensively generated samples cover the support of the reference distribution.
  • Diversity captures motion variability.

Geometric fidelity

  • Chamfer compares root-aligned SMPL-X vertex clouds through average nearest-neighbor distance.
  • Hausdorff focuses on the largest unmatched surface discrepancy.
  • MJD measures root-aligned mean joint distance for body and hands.
  • Dice measures overlap between projected motion occupancy maps.
  • PCK evaluates joint correctness under a distance threshold.

Motion quality

  • LDLJREL compares normalized movement smoothness; values close to zero indicate smoothness similar to the reference.
  • Foot Contact compares stable ground-contact rates and targets lower-body plausibility and sliding artifacts.

Speech-motion alignment

  • Beat Consistency (BC) measures temporal correspondence between detected motion beats and acoustic beats.
  • BC is prosody-aware but semantics-agnostic: it measures when gestures occur, not whether they communicate the spoken meaning.

Semantic appropriateness

  • SRGR upweights pose correctness at semantically relevant frames.
  • Semantic Score (SC) compares speech and motion embeddings with cosine similarity.
  • SGP measures whether semantic gesture intervals avoid collapsing into generic beat-like motion.

Why reference-pose metrics are limited

Co-speech gesture realization is one-to-many. The same communicative intent may be expressed with a different hand path, orientation, scale, or even a different valid gesture. Exact geometric agreement can therefore underestimate perceptually appropriate motion.

Why distribution metrics matter

FGD, Density, and Coverage characterize whether generated motion occupies plausible regions of motion space. In our results, these distributional measures are often more predictive of human ratings than direct pose-space distances.

Why semantic metrics need interpretation

SRGR and SC provide useful semantic signals, but neither directly resolves the many-to-many relation between meaning and motion. SGP is designed as a complementary event-level measure rather than a replacement for all other metrics.

03 · Perceptual validation

Human evaluation is split into visual and speech-aware conditions.

Participants complete only one condition, so motion-centric judgments can be separated from judgments that require listening to speech. All criteria use a seven-point scale, with higher ratings indicating better perceived quality.

🔇 Muted condition · 49 valid participants

Human-likenessCould a real person perform this motion?
Motion diversityDoes the animation contain varied rather than repetitive movement?
Absence of animation errorsIs the motion free from jumps, freezing, interpenetration, or strange behavior?

🔊 Audio condition · 52 valid participants

Speech timingDo gestures occur at appropriate moments relative to rhythm, pauses, and emphasis?
Content matchDo gestures meaningfully correspond to the spoken content?

Stimulus construction

The study compares the four generated systems with BEAT2 motion-capture ground truth. All conditions use the same avatar, camera, crop, renderer, and 10-second temporal window.

Semantic enrichment

Windows are selected to contain many iconic, metaphoric, and deictic gestures. The selected stimuli contain 80 annotated semantic events: 44 iconic, 18 metaphoric, and 18 deictic.

Quality control

Each study includes practice trials and 16 attention checks. Videos must be watched in full before ratings are submitted. Participants failing two or more checks are excluded.

103 participants were recruited; 101 remained after quality control. Two participants were excluded from the muted study, leaving 49 muted and 52 audio participants. Repeated sequence observations were pooled before aggregation, producing 75 model-sequence observations across all five motion sources and 60 generated-motion observations after excluding ground truth.
04 · Semantics-aware evaluation

Fine-grained semantic annotation with Gemini 2.5 Pro.

Existing BEAT annotations specify semantic intervals and coarse categories, but do not fully describe how a gesture is physically realized or what it means in context. We enrich those annotations instead of replacing them.

Event-level annotation

For each semantic event, the rendered motion segment and target word are provided to Gemini 2.5 Pro. The model is instructed to analyze only the specified interval and ground its output in visible motion rather than assuming a semantic gesture is present.

Handedness
Hand shape
Orientation
Location
Movement
Gesture description Contextual meaning Gesture type Confidence

Semantic Gesture Preservation (SGP)

SGP focuses on whether a semantic interval is preserved as a non-generic gesture rather than collapsing into a beat. For each event, Gemini classifies the observed gesture as iconic, metaphoric, deictic, beat, none, or uncertain.

SGP = 1 − (# semantic events classified as beat / # valid semantic events)

SGP = 1 means none of the evaluated semantic events are classified as beats. The score is averaged at the clip level, so each clip contributes equally regardless of event count.

Manual validation

A randomly sampled subset of Gemini outputs is manually checked against the rendered motion. An annotation is counted as correct when its visible gesture description captures the principal motion and its semantic interpretation is compatible with the manually assessed communicative function. Under this criterion, 75% of the evaluated LLM-generated annotations are judged correct.

Interpretation boundary

SGP is not semantic ground truth and does not verify exact meaning preservation or exact gesture category. It is specifically a measure of semantic-gesture preservation against beat-gesture collapse, designed to relax strict one-to-one pose matching.

05 · Objective benchmark

No model dominates every metric family.

The benchmark results show that distributional similarity, geometry, timing, physical plausibility, and semantic preservation do not improve together automatically. This is exactly why a complementary metric suite is needed.

MethodFGD ↓BC ↑Div. ↑Dens. ↑Cov. ↑Dice ↑ LDLJREL →0SRGR ↑PCK ↑MJD ↓Chamfer ↓Hausdorff ↓ Foot Contact ↓SC ↑SGP ↑
EMAGE5.5230.692880.89720.072750.6407 0.05300.054940.33090.24100.013750.29080.50480.2480.4372
SemConFlow2.2450.7801200.015590.066470.6793 -0.42940.058680.35520.20070.011120.25300.59320.3140.4205
SemTalk4.2660.7271160.11230.066730.6241 -0.60200.050590.30370.23030.139760.46630.40400.2680.4014
GestureLSM4.2680.5251120.0092390.0098170.6573 1.53860.045360.27150.23910.017970.29090.88830.2480.2622

SemConFlow

Achieves the strongest performance on most distributional, geometric, and speech-alignment metrics, including FGD, BC, Diversity, Dice, SRGR, PCK, MJD, Chamfer, Hausdorff, and SC.

EMAGE

Obtains the best Density and Coverage, has LDLJREL closest to zero, and achieves the highest SGP score, indicating fewer semantic intervals collapse into generic beats.

SemTalk & GestureLSM

SemTalk achieves the best Foot Contact result, while GestureLSM remains competitive on diversity and geometric fidelity but does not lead an individual metric in the benchmark table.

Objective metrics are computed over 265 shared test sequences, except SGP, which is evaluated on the 15 perceptual-study sequences with Gemini-based semantic annotations.

06 · Subjective results

Ground truth remains clearly ahead — especially on perceptual quality and speech appropriateness.

BEAT2 reference motion receives the highest average rating on all five human-study dimensions. The gap is smaller for motion diversity, suggesting that current systems can approach reference-level variability more easily than they reproduce naturalness and speech-conditioned appropriateness.

SemConFlow leads the generated systems.

Among generated methods, SemConFlow achieves the highest mean subjective score on all five perceptual dimensions.

Muted and audio rankings differ.

In the muted study, SemTalk generally ranks second. In the audio study, EMAGE ranks second for speech timing and content match.

Visual quality is not speech appropriateness.

A model that looks comparatively good without audio does not necessarily keep the same advantage when timing and semantic correspondence are judged.

GestureLSM receives the lowest subjective means.

GestureLSM consistently receives the lowest mean ratings among the generated systems in the reported perceptual results.

07 · Objective–subjective correlation

Which automatic metrics actually track human perception?

The answer depends strongly on the perceptual target. No individual metric consistently explains all five human judgments, and Semantic Score (SC) shows no statistically meaningful association with any of the five subjective dimensions.

Human-likeness

LDLJREL−.61
Coverage.55

Motion diversity

Coverage.52
Density.50

Absence of animation errors

LDLJREL−.62
Coverage.54

Speech timing

Coverage.73
Density.68

Content match

Coverage.70
Density.66

Muted judgments

LDLJREL is the strongest correlate for both human-likeness (ρ = −0.61) and absence of animation errors (ρ = −0.62). Coverage and Density are most informative for motion diversity.

Speech-aware judgments

Coverage reaches ρ = 0.73 for speech timing and ρ = 0.70 for content match. Density and FGD show similarly strong associations with both speech-dependent targets.

Geometry is comparatively weak

MJD and Hausdorff show little correspondence with subjective dimensions, reinforcing that perceptually appropriate gestures need not reproduce the exact reference geometry.

Beat Consistency behaves counter-intuitively.

BC is negatively correlated with all five perceptual dimensions, indicating that stronger beat alignment alone does not guarantee more natural or communicatively appropriate gestures.

SGP is selectively speech-aware.

SGP shows the intended selectivity toward speech-aware judgments, correlating with both speech timing and content match at approximately ρ = 0.39 while remaining weak or non-significant for muted visual dimensions. SC, in contrast, shows no statistically meaningful correlation with any of the five subjective dimensions.

08 · Composite metrics

Can combinations of metrics predict perception better than single metrics?

For each perceptual target, a separate ordinary least-squares regressor combines 13 objective predictors. Chamfer Distance is excluded from the joint model because of its strong redundancy with Hausdorff Distance, while SC is retained as an independent semantics-aware predictor. Models are evaluated with leave-one-out cross-validation over the 60 generated model-sequence observations, with feature standardization performed inside each fold.

Target-specific fitting

Separate regressors are trained for human-likeness, motion diversity, absence of animation errors, speech timing, and content match. This avoids assuming that the same metric mixture should explain every perceptual dimension.

Evaluation

Held-out predictions are compared with human ratings using Spearman correlation, mean squared error, and mean absolute error.

Interpretation

All five composites improve over their strongest individual metric. The largest gains occur for absence of animation errors and content match, while the benefits for motion diversity and speech timing are smaller, showing that metric combination remains strongly target-dependent.

Target compositeMSE ↓MAE ↓Composite ρ ↑Best individualΔ|ρ| ↑
Human-likeness0.28460.41400.6568LDLJREL+0.0468
Motion diversity0.20350.36100.5375Coverage+0.0117
Absence of animation errors0.15600.29840.7065LDLJREL+0.0819
Speech timing0.26710.43490.7439Coverage+0.0246
Content match0.36650.48440.7824Coverage+0.0805

Composite predictors are evaluated as dimension-specific automated proxies for human judgments, using out-of-fold predictions on the original seven-point rating scale rather than deriving an additional model-level ranking.

09 · Limitations

What this benchmark does — and does not — establish.

Perceptual sample size

The human correlation analysis is based on 15 unique speech sequences across four generated systems, yielding 60 generated model-sequence observations rather than 60 fully independent utterances.

Semantic enrichment

The selected windows deliberately contain many semantic gestures. This improves sensitivity to semantic differences but may overrepresent semantically dense speech relative to the broader BEAT2 distribution.

Online study conditions

Listening environment, distraction, fatigue, and presentation-order effects cannot be fully controlled in an online perceptual study.

Gemini annotations

Gemini 2.5 Pro provides scalable enrichment, but its outputs should not be treated as semantic ground truth.

SGP scope

SGP measures preservation against beat-gesture collapse; it does not directly verify that the exact intended meaning or semantic gesture category is preserved.

Generalization

The observed objective–subjective relationships are specific to the evaluated models, sequences, and metric implementations. Future work should extend the benchmark to broader datasets, models, and human studies.

Takeaway
Reliable progress in holistic co-speech gesture generation requires multiple perceptually validated metrics, not a single aggregate score.

Distributional similarity, geometric fidelity, temporal alignment, physical plausibility, and semantic preservation each capture a different slice of gesture quality. The benchmark therefore treats evaluation as a multidimensional problem: standardize the protocol, validate metrics against human perception, and use semantics-aware measures as complementary evidence rather than universal substitutes.