Expressive talking-face generation

Make the face match the feeling.

VAExpress turns one portrait + one audio clip into a natural talking-face video. It goes beyond lip sync by generating facial expressions that follow the emotion of the voice.

  • Emotion from audio
  • Continuous emotion control
  • Identity-aware generation
Audio-driven results across different emotions. Turn on sound to view the full example.
Goal

Generate a talking face that is not only synchronized, but also emotionally expressive.

Method

Continuous Valence-Arousal guidance with spatial-temporal decoupled training.

Result

Readable expressions, controllable intensity, and smooth emotion transitions.

Generated results

Judge it in motion.

Three independent windows show method comparison, audio-driven emotion, and continuous emotion control.

Example 01
Example 02
Method comparison

More expressive than lip-sync-first methods

On the same portrait and audio, VAExpress produces clearer motion in the eyes, brows, cheeks, and overall facial expression.

Example 01
Example 02
Audio-driven emotion

The voice can direct the expression

When no manual control is supplied, emotion cues are estimated from speech and used to guide the facial performance.

Emotion control
Continuous emotion control

Adjust intensity and transition smoothly

A continuous emotion sequence can raise expression intensity gradually or move between affective states without abrupt category changes.

Audio-aligned

Expressions follow the affective state of speech.

Fine-grained

Intensity can change continuously over time.

Temporally coherent

Transitions stay smooth without flattening expression.

How it works

Three inputs, one expressive video.

VAExpress model pipeline and two-stage training strategy
Technical view

Overall pipeline and two-stage training

Reference identity, speech, and continuous VA emotion signals enter separate conditioning branches before being combined by the diffusion generator. The right side shows the image-first, video-second training strategy.

Continuous guidance

Valence controls feeling. Arousal controls energy.

Instead of choosing only a fixed label such as “happy” or “sad,” VAExpress uses a continuous two-dimensional emotion signal. This makes subtle intensity changes and smooth transitions possible.

Generated facial expressions distributed across the Valence-Arousal emotion space

Decoupled training

Learn expression first, then learn smooth motion.

Training spatial and temporal modules together can over-smooth facial movement. VAExpress first learns accurate emotion-to-expression mappings on images, then adds temporal training for coherent video.

  1. Stage 1Static expression learning
  2. Stage 2Temporal adaptation
Audio-aligned facial expressions, fine-grained intensity adjustment, and emotion transition
Spatial-temporal decoupled training visualization

Clear expression response, then smooth temporal change.

The top row shows expressions aligned with speech. The bottom row shows that continuous VA signals can adjust intensity and drive an emotion transition. The two-stage training keeps this expressive response before adding temporal coherence.

Why decouple training

Temporal modules can suppress expression response.

The orange curves with temporal modules are flatter and react more slowly to changing arousal. The blue curves without them follow the control signal more closely. This motivates learning expressive mappings first, then adding temporal coherence.

Analysis of spatial suppression and temporal smoothing caused by temporal modules

The practical effect is visible in the clips above. Controlled evaluation also shows stronger expression accuracy and video quality on MEAD and RAVDESS.

View the full evaluation in the technical report ↗

VAExpress

One portrait. One voice.
A more expressive result.

Open technical report ↗