Introduction

Imagine watching a virtual talk show where the host delivers engaging dialogue complemented by expressive gestures, natural body movements, and precise movement paths. The host walks across the stage following a scripted trajectory, uses hand gestures to emphasize points based on their speech, and shifts posture in response to both the conversation’s flow and predefined text instructions—all occurring in perfect harmony. This level of realism transforms the viewing experience, making interactions feel genuine and immersive. Achieving such lifelike behavior is no small feat, yet it is essential for enhancing user engagement in applications ranging from virtual reality to interactive gaming and beyond.

Driving a 3D avatar to perform such lifelike motions involves managing multiple control signals, such as text descriptions, speech audio, and trajectory data. Multi-modal signals may be provided concurrently, for instance, a text prompt like “a man is walking” alongside a speech audio clip. However, most prior works primarily focus on single-modality control, such as text-to-motion [20, 59] or speech-to-gesture [18, 71]. Recent studies [76, 82, 83] have explored designing unified models capable of addressing multi-modal signals by leveraging datasets from different generation tasks. Nevertheless, these models typically process only one modality at a time, combining motions conditioned on different inputs in a limited and sequential manner when multiple control signals are present.

Examples of Multi-Modal Controlled Motion Generation with text descriptions, speech audio, and trajectory data

Figure 1. Examples of Multi-Modal Controlled Motion Generation. Given multiple control signals of different modalities—including text descriptions, speech audio, and trajectory data—our MOCO framework generates realistic and coherent holistic body motion. This includes both body movements and detailed features such as facial expressions and hand gestures, all closely aligned with the provided conditions. To clearly illustrate this, we highlight two clips with temporal zoom, showcasing the natural integration of speech gestures and lower-body movements in our generated motions. Additional visual results are provided in Figure 4 and the supplementary materials.

The primary challenge in achieving simultaneous multi-modal control of motion generation is the scarcity of aligned multi-modal data. While collecting additional multi-modal data could help, it requires significant resources. In addition, the activity regions in speech-to-gesture datasets are often limited, making it hard to train models that can generate trajectory-controlled speech gestures. Some efforts address this issue by combining the predictions of text-driven model and audio-driven model through weighted sum [12, 70]. [37] suggested using speech scripts as pseudo text labels to create aligned text-audio-motion datasets, replacing scripts with movement descriptions during inference. However, these approaches face inherent limitations: for weighted sum, the predominance of standing poses in speech-to-gesture datasets creates imbalanced model corrections; for pseudo-labels, the use of speech scripts as training labels limits generalization to diverse motion descriptions during inference. To overcome these challenges, we propose a novel diffusion-based framework, named Multi-MOdal COherent Motion Synthesis (MOCO). MOCO exploits a natural spatial decomposition: speech audio primarily drives upper-body dynamics—gestures and facial expressions—while text descriptions mainly influence lower-body movements such as walking and stance shifts (see Appendix A for supporting statistics). To achieve this, MOCO is first trained on multiple datasets, ensuring that the model can independently generate motions conditioned on either text or speech inputs. At each denoising step, the model generates motions for each modality separately from the input noise and assembles the body parts according to predefined spatial rules, i.e. combining audio-driven upper-body motion with text-driven lower-body motion. This combined motion is then diffused and used as the input noise for the next denoising step. The separation ensures that each body part’s motion is highly aligned with its corresponding input condition, while the iterative process conditions each generation step on the current state of the combined motion. This allows each modality to refine its contribution within the context of the overall movement. Consequently, with each iteration, the motions generated for different body parts become increasingly harmonized, resulting in natural movements that exhibit coherent and synchronized behaviors. Furthermore, this decoupled generation process enables our framework to incorporate trajectory control into co-speech motion generation. We can leverage trajectory data to generate text-driven motion and combine it with audio-driven motion, producing speech gestures that closely align with the given trajectory.

To facilitate the evaluation of this novel task, we develop a multi-modal benchmark comprising 1,000 test clips which are generated from 40 fundamental text descriptions of body movements (e.g., “walk forwards” and “step back and sit down”) and 694 audio clips from eight different speakers. Each test clip integrates two text prompts describing a movement with two speech audio clips. We rigorously evaluate our approach against baseline methods using both text-to-motion and speech-to-gesture metrics. Experimental results demonstrate that our method outperforms existing baselines, advancing the field of multi-modal controlled motion generation for 3D avatars.

Multi-Modal Conditioned Motion Generation

In recent years, human motion generation has received significant attention, largely driven by advancements in dataset collection. Various scenarios have been explored depending on the input conditions, including action labels [21], text descriptions [6, 14, 20, 23, 26, 27, 35, 59, 63, 75, 78], speech audio [18, 31, 40, 42, 44, 50, 53, 71, 80], music [34, 55, 60], scene context [24, 47], spatial signal [58, 66], and even the motion of another person [11, 22, 46, 48, 57, 62]. Beyond single-modality control, research efforts have been extended to handle integrating multi-modal signals. For instance, some works considered speaker identity, speech audio, and transcripts to generate conversational gestures [2, 3, 68, 73], while other works jointly leverage text and scene inputs for motion generation [10, 15, 64, 72]. Moreover, some works focus on integrating various datasets to train unified motion models that enhance scalability and applicability across multiple scenarios [13, 28, 76, 81, 83].

A critical limitation persists, however: current models depend heavily on aligned multi-modal training data, limiting their ability to handle novel input combinations (e.g., text with audio, audio with trajectory) unseen in training. To address this, FreeTalker [70] and SynTalker [12]—the latter concurrent with the first version of this paper [45]—propose combining predictions from text-driven and audio-driven models through weighted sum. Liu et al. [37] suggested using speech scripts as pseudo text labels (e.g., “A male speaker is saying: ‘I am shocked by what you have done.’ ”) to create aligned text-audio-motion datasets for training and replace these scripts with movement descriptions during inference. However, the weighted sum method poses challenges, as the speech-driven model exerts significantly stronger correction forces than the text-driven model. This imbalance often generates motions that overly favor speech conditions while suppressing text conditions (Appendix E). Regarding pseudo labels, the reliance on speech transcripts as supervisory signals creates a domain gap between the transcript distribution and the motion-description distribution, thereby constraining generalization to diverse motion descriptions.

Diffusion-based Methods in Motion Generation

The diffusion model has been recognized as one of the most advanced generative paradigms and has also gained significant attraction in human motion generation. Its applications can be categorized into two primary streams according to the space where the denoising process is implemented. One representative method in the first stream is the Human Motion Diffusion Model (MDM) [59], which operates directly in the original motion space to perform denoising [1, 29, 32, 36, 84]. This approach offers several advantages, including high editability and strong controllability, enabling operations such as concatenation and combination within the original motion space during the denoising process. The second stream is exemplified by the Motion Latent Diffusion model (MLD) [14], which conducts denoising in the latent space of a Variational Autoencoder (VAE) [8, 16, 52].

Utilizing the VAE latent space improves efficiency in representing complex motion data and significantly reduces the computational load required for diffusion processes, resulting in faster inference speeds. In this work, we select MDM as the foundational model due to its superior editability [5, 49, 54].

Compositional Motion Generation

A complementary line of work generates complex motion by composing multiple actions or sub-motions. Spatial composition assembles motions of different body parts: SINC [5] combines multiple text-driven actions through body-part masks in a VAE-based, offline manner. Temporal composition instead sequences multiple actions over time and focuses on seamless transitions between them, as in TEACH [4], T2LM [33], and FlowMDM [9]. STMC [49] unifies both axes, bringing per-denoising-step body-part composition into the diffusion framework with a DiffCollage-style [79] mechanism for temporally overlapping actions. All these methods compose homogeneous, text-only conditions; we instead build on the per-step composition paradigm and extend it to heterogeneous audio, text, and trajectory control, placing MOCO at the intersection of compositional and multi-modal motion generation.

Method

In this section, we begin with a brief introduction to the Motion Diffusion Model (MDM) [59], which serves as our foundational model (Sec. 3.1). Then, we describe our framework’s data representations and important model components (Sec. 3.2). Next, we explain our multi-modal decoupled denoising strategy for holistic motion generation (Sec. 3.3). Finally, we address a more complex scenario where the trajectory condition is included and utilize a Large Language Model for motion planning (Sec. 3.4). An overview of our framework can be found in Figure 2.

Overview of MOCO, showing decoupled denoising and combination of detailed facial and hand movements with body motion

Figure 2. Overview of MOCO. At each denoising step tt, the input conditions and noisy data are fed into their respective denoisers to predict clean motion, which is then diffused for the next iteration. Specifically, the upper-body motion conditioned on speech audio and the lower-body motion conditioned on text description are combined to form the overall body motion. The blue arrows in the figure highlight two key points. One indicates that the denoising process of v0v_0 is completed before body motion denoising. The other shows that after the denoising process, the detailed facial and hand movements, and the combined body motion are integrated together to produce the final holistic motion.

Preliminary: Motion Diffusion Model

MDM is a diffusion-based motion generation model [59]. The same as other diffusion models, it regards diffusion as a Markov noising process (xt)t=0T(x_t)_{t=0}^{T} that starts from a training sample x0x_0. As the time step tt increases, the distribution of xTx_T approaches a standard normal distribution. Besides, MDM models the conditional distribution p(x0∣c)p(x_0 \mid c) by reversing the diffusion process through iterative denoising on xTx_T. To achieve this, it minimizes a loss function that penalizes the difference between the original and denoised samples. Moreover, it performs sampling from p(x0∣c)p(x_0 \mid c) iteratively, predicting x0x_0 at each timestep tt and computing xt−1x_{t-1} until tt reaches 0.

MDM adopts the classifier-free guidance [25] to adjust its adherence to the conditioning signal cc. Specially, its denoiser GG is trained on both conditioned and unconditioned data by randomly setting cc as ∅\varnothing, which allows G(xt,t,∅)G(x_t, t, \varnothing) to approximate the unconditional data distribution. During sampling, it adjusts the strength of cc using a scaling factor ss:

Gs(xt,t,c)=G(xt,t,∅)+s⋅(G(xt,t,c)−G(xt,t,∅)),(1)G^s\left(x_t,t,c\right) = G\left(x_t,t,\varnothing\right) + s \cdot\left(G\left(x_t,t,c\right) - G\left(x_t,t,\varnothing\right)\right), \tag*{(1)}

where GsG^s denotes the denoiser with the classifier-free guidance.

Data Representation and Model Architecture

Data Representation. Our framework incorporates four data modalities: motion, text, audio, and trajectory. The motion data is represented as m={mn}n=1N∈RN×491m = \{m^n\}_{n=1}^{N} \in\mathbb{R}^{N \times491}, where NN is the number of frames. Specifically, the motion data for each frame is denoted as mn={bn,dn}m^n = \{b^n, d^n\}, with b∈R205b \in\mathbb{R}^{205} representing the body pose [49], and d∈R286d \in\mathbb{R}^{286} capturing detailed facial expression and hand movements. The text embeddings are encoded using the CLIP model [51] and are denoted as ctext∈R512c_{text} \in\mathbb{R}^{512}. Audio features are extracted via the Wav2Vec2 model [7] and represented as caudio∈RN×768c_{audio} \in\mathbb{R}^{N \times768}. Finally, the trajectory data is encoded as ctraj∈RN×2c_{traj} \in\mathbb{R}^{N \times2}, representing the position on the XY-plane for each frame.

Model Design. Our framework includes four transformer-based denoisers: one for text-to-motion (T2M), one for speech-to-gesture (S2G), one for trajectoryto-velocity (T2V), and one for speech-to-details (S2D) that synthesizes facial expressions and hand poses:

b^0=GT2M(bt,t,ctext),b^0=GS2G(bt,t,caudio)(2)\hat{b}_0 = G_{\mathrm{T2M}}(b_t, t, c_{\mathrm{text}}), \qquad\hat{b}_0 = G_{\mathrm{S2G}}(b_t, t, c_{\mathrm{audio}}) \tag*{(2)}
v^0=GT2V(vt,t,ctraj),d^0=GS2D(dt,t,caudio)(3)\hat{v}_0 = G_{\mathrm{T2V}}(v_t, t, c_{\mathrm{traj}}), \qquad\hat{d}_0 = G_{\mathrm{S2D}}(d_t, t, c_{\mathrm{audio}}) \tag*{(3)}

We denote the sampling with classifier-free guidance for each denoiser as GusG_u^s, where u∈{T2M,S2G,T2V,S2D}u \in\{\mathrm{T2M},\mathrm{S2G},\mathrm{T2V},\mathrm{S2D}\}.

For the T2M denoiser conditioned on a text embedding ctextc_{\mathrm{text}}, we follow prior works [14, 59] by treating ctextc_{\mathrm{text}} as a token and applying self-attention to incorporate the semantic information in ctextc_{\mathrm{text}} into the motion generation process. In contrast, the other denoisers are conditioned on sequential data; therefore, we utilize cross-attention to model the relationships between the input sequences and the generated motion. Additionally, the T2M and T2V denoisers are trained on HumanML3D [20], while the S2G and S2D denoisers are trained on the BEAT2 [40] dataset. All denoisers adhere to the objective function and diffusion paradigm described in Sec. 3.1. The computational complexity of each component is reported in Appendix F.

Multi-Modal Decoupled Denoising

In this section, we introduce our multi-modal decoupled denoising strategy for generating motion in scenarios where text and speech audio are provided synchronously—that is, within the same time interval—as shown in Figure 3 (a).

Examples of synchronous and asynchronous conditions

Figure 3. Examples of synchronous and asynchronous conditions. Synchronous conditions occur when all condition signals are provided within the same time interval. In contrast, asynchronous conditions involve multiple conditions, each corresponding to different time intervals.

Our key observation is that the speech audio naturally guides upper-body gestures (including head and arm poses), while text descriptions mainly influence lower-body movements (including spine and leg poses) like walking or shifting stance. In a diffusion context, this implies that the joint conditional probability p(bt−1∣ctext,caudio,bt)p(b_{t-1} \mid c_{\mathrm{text}}, c_{\mathrm{audio}}, b_t) can be approximated by the product of two probabilities: p(bt−1,lower∣ctext,bt)p(b_{t-1,\mathrm{lower}} \mid c_{\mathrm{text}}, b_t) and p(bt−1,upper∣caudio,bt)p(b_{t-1,\mathrm{upper}} \mid c_{\mathrm{audio}}, b_t). Mathematically, this can be represented as follows:

p(bt−1∣ctext,caudio,bt)≈p(bt−1,lower∣ctext,bt)⋅p(bt−1,upper∣caudio,bt).(4)p(b_{t-1} \mid c_{\mathrm{text}}, c_{\mathrm{audio}}, b_t) \approx p(b_{t-1,\mathrm{lower}} \mid c_{\mathrm{text}}, b_t) \cdot p(b_{t-1,\mathrm{upper}} \mid c_{\mathrm{audio}}, b_t). \tag*{(4)}

Here, btb_t represents the motion at denoising step tt, which is composed of the upper-body motion bt,upperb_{t,\mathrm{upper}} and the lower-body motion bt,lowerb_{t,\mathrm{lower}}. A detailed derivation of (4) is provided in Appendix B.

Based on this observation, we can decouple the multi-modal controlled denoising process into distinct streams. The text-driven stream p(bt−1∣ctext,bt)p(b_{t-1} \mid c_{\mathrm{text}}, b_t) can be formulated as follows:

b^0text=GT2Ms(bt,t,ctext),bt−1text=αt−1 b^0text+1−αt−1 ϵ(5)\hat{b}^{\mathrm{text}}_0 = G^s_{\mathrm{T2M}}(b_t, t, c_{\mathrm{text}}), \qquad b^{\mathrm{text}}_{t-1} = \sqrt{\alpha_{t-1}}\,\hat{b}^{\mathrm{text}}_0 + \sqrt{1-\alpha_{t-1}}\,\epsilon \tag*{(5)}

where αt=∏s=1t(1−βs)\alpha_t = \prod_{s=1}^{t}(1-\beta_s) represents the cumulative product of (1−βs)(1-\beta_s) up to timestep tt, and βs\beta_s is the variance schedule controlling the amount of noise added at each timestep. The variable ϵ∼N(0,I)\epsilon\sim\mathcal{N}(0,\mathbf{I}) is Gaussian noise sampled from a standard normal distribution [59]. Similarly, the audio-driven stream p(bt−1∣caudio,bt)p(b_{t-1} \mid c_{\mathrm{audio}}, b_t) is defined as:

b^0audio=GS2Gs(bt,t,caudio),bt−1audio=αt−1 b^0audio+1−αt−1 ϵ.(6)\hat{b}^{\mathrm{audio}}_0 = G^s_{\mathrm{S2G}}(b_t, t, c_{\mathrm{audio}}), \qquad b^{\mathrm{audio}}_{t-1} = \sqrt{\alpha_{t-1}}\,\hat{b}^{\mathrm{audio}}_0 + \sqrt{1-\alpha_{t-1}}\,\epsilon. \tag*{(6)}

Finally, the body motion bt−1b_{t-1} at step tt is computed as follows:

bt−1=I⊙bt−1text+(1−I)⊙bt−1audio,(7)b_{t-1} = I \odot b^{\mathrm{text}}_{t-1} + (1-I) \odot b^{\mathrm{audio}}_{t-1}, \tag*{(7)}

where I∈R205I \in\mathbb{R}^{205} is a binary vector with entries set to 1 for the lower body and 0 for the upper body; ⊙\odot denotes element-wise multiplication.

The decoupled denoising allows each body part’s motion to be precisely guided by its corresponding input condition, ensuring high fidelity to the control signals. Furthermore, each stream adjusts its output based on the current overall motion, promoting coordination among the body parts. As this process continues, the motions generated for different body parts become increasingly synchronized, resulting in natural and coherent full-body movements.

Trajectory and Motion Planning

The above subsection handles the multi-modal decoupled denoising for synchronous conditions. We now extend our MOCO framework to tackle more complex, challenging yet more realistic scenarios, such as incorporating trajectory control and leveraging Large Language Model (LLM) for motion planning.

Trajectory Control. Following the approach of Petrovich et al. [49], we represent the global transition of body pose using the velocity vector v=[r˙x,r˙y,θ˙]v = [\dot{r}_x,\dot{r}_y,\dot{\theta}], where r˙x\dot{r}_x and r˙y\dot{r}_y are the linear velocities of the pelvis in the xx and yy directions, respectively, and θ˙\dot{\theta} is the angular velocity about the body’s vertical (ZZ) axis. Given the trajectory data ctrajc_{\mathrm{traj}}, we first predict v^0\hat{v}_0 using (3). To enhance prediction accuracy, we incorporate loss guidance into our method. During each denoising step for predicting the velocity vector, we compute v^0\hat{v}_0 using (3) and apply loss guidance as follows:

Lguidance=FK(v^0)−ctraj,(8)L_{\mathrm{guidance}} = FK(\hat{v}_0) - c_{\mathrm{traj}}, \tag*{(8)}

where FKFK represents the differentiable Forward Kinematics function [65] that converts linear and angular velocities into the trajectory. We follow the methodology of InterControl [65] and optimize LguidanceL_{\mathrm{guidance}} with respect to v^0\hat{v}_0 using the second-order LBFGS optimizer [38]. This optimization ensures that the predicted global transitions closely match the provided trajectory data.

Once v^0\hat{v}_0 is predicted, we substitute the velocity component in b^0\hat{b}_0 with v^0\hat{v}_0 during each iteration of its generation. This substitution guides the generation process to adapt the remaining elements of b^0\hat{b}_0 to align with v^0\hat{v}_0, thereby ensuring consistency with the provided trajectory data.

Motion Planning. In real-world applications, models could face audio conditions paired with complex text descriptions, which can challenge their ability to handle intricate semantics. One approach to address this challenge is decomposing complex textual descriptions into basic motion units and generating them separately [56]. However, this decomposition often introduces more complex conditions—such as asynchronous conditions, as illustrated in Figure 3.

To address these challenges, we propose an LLM-based motion planning pipeline. Our pipeline begins by using an LLM to generate a “motion timeline” from the input text and audio conditions. Particularly, we carefully design a prompt that instructs the LLM to decompose complex text into elementary motion units, insert appropriate pose transitions, and specify the start and end times and the body parts involved for each condition. More details are provided in Appendix C. For example, when given the text condition “kneel down then walk forward” along with two audio clips, the LLM produces a motion timeline comprising multiple intervals, as detailed in Table 1.

ConditionStart TimeEnd TimeBody Part
c1=c_1= kneel downf1sf_1^sf1ef_1^espine, legs
c2=c_2= stand up from groundf2sf_2^sf2ef_2^espine, legs
c3=c_3= run forwardf3sf_3^sf3ef_3^espine, legs
c4=c_4= audio clip 1f4sf_4^sf4ef_4^ehead, arms
c5=c_5= audio clip 2f5sf_5^sf5ef_5^ehead, arms

Table 1. Example of an LLM-generated motion timeline.

ConditionStart TimeEnd TimeBody Part
sit down on a chair0.05.0legs, spine
stand up from chair5.07.0legs, spine
turn around7.011.5legs, spine
walk forward11.517.5legs, spine
5_stewart_0_1_1,$785440$9113285.012.85head, arms
5_stewart_0_1_1,$918560$105008017.524.3head, arms

Table 1.

ConditionStart TimeEnd TimeBody Part
a man walks then waves hands0.08.0legs, spine
speak:1_wayne_0_103_103,$9248$998082.5788.238left arm, right arm, head
speak:1_wayne_0_103_103,$107552$1960649.72215.254left arm, right arm, head

Table 1.

Notably, the transition condition c2c_2 is automatically generated by the LLM, which ensures a smooth stance changes and avoids abrupt movements. Users, however, have the flexibility to customize the timeline. We then adopt the strategy of STMC [49], generating body masks according to a predefined timeline. At each denoising step tt, the body pose across the entire timeline is generated as follows:

b^0=∑j=1JIj⊙Gjs(bt,fjs:fje,t,cj),(9)\hat{b}_0 = \sum_{j=1}^{J} I_j \odot G_j^s\left(b_{t,f_j^s:f_j^e}, t, c_j\right), \tag*{(9)}

where IjI_j represents a binary mask corresponding to the motion generated by the jj-th condition, and Gjs∈{GT2Ms,GS2Gs}G_j^s \in\{G_{\mathrm{T2M}}^s, G_{\mathrm{S2G}}^s\} denotes the denoiser utilized for the jj-th condition. The operator ⊙\odot stands for element-wise multiplication.

Similarly, the facial and hand movements at step tt are computed by:

d^0=∑j=1JGS2D(dt,fjs:fje,t,cj).(10)\hat{d}_0 = \sum_{j=1}^{J} G_{\mathrm{S2D}}\left(d_{t,f^s_j:f^e_j}, t, c_j\right). \tag*{(10)}

Since the text-to-motion dataset, HumanML3D, lacks facial and hand movement data, the denoiser GS2DG_{\mathrm{S2D}} is only trained on the speech-to-gesture dataset. Therefore, in intervals without audio clips, we maintain the model’s temporal continuity by setting the conditional input cjc_j to the empty condition ∅\varnothing. Finally, we employ DiffCollage [79] to ensure smoother transitions at the interval boundaries.

Experiments

Datasets

Task-Specific Datasets. The HumanML3D dataset is a large Text-to-Motion dataset created by amalgamating motion sequences from the Human-Act12 and AMASS datasets [20]. It consists of 14,616 motions and 44,970 descriptions composed of 5,371 distinct words, totaling 28.59 hours of motion data. To align the data representation—specifically, to use SMPL-X parameters for representing joint rotations—we utilize only the AMASS portion of HumanML3D because it has an official SMPL-X version. The BEAT2 dataset is a large-scale Speech-to-Gesture dataset specifically designed for research in speech-to-gesture generation [40]. It contains synchronized recordings of speech audio and corresponding 3D motion capture data of human gestures, totaling 60 hours of data from 25 speakers. In addition to audio and motion data, the dataset includes valuable annotations such as text transcriptions and emotional states. Here, we select speakers with IDs lower than 10, resulting in a total of 10.15 hours of training data.

Multi-Modal Benchmark. We create a multi-modal benchmark of 1,000 test clips to evaluate our proposed task effectively. Each test clip is automatically constructed and contains two text descriptions and two audio clips. To create these clips, we manually collect 40 texts focusing on lower-body movements that commonly occur during speech delivery or conversation. We then split the audio from the BEAT2 test set into clips using a Voice Activity Detector (VAD). To serve as ground truth for computing evaluation metrics (Sec. 4.2), we selected motion samples from AMASS and BEAT2 corresponding to each text and audio clip. This structured construction enables controlled and repeatable comparisons under concurrent multi-modal conditions and isolates the challenge of resolving conflicting cross-modal signals. Accordingly, the benchmark is intended to probe compositional consistency rather than broad coverage of open-ended semantics. More details about the benchmark are provided in Appendix H.

Metrics

We evaluate our method using two categories of metrics: text-to-motion and speech-to-gesture [20, 40, 49]. For text-to-motion, FID+FID+ assesses realism by measuring the distribution difference between real and generated motions using five random 5-second clips per test sample. R1R1 evaluates alignment by recording the frequency of correct text prompts appearing in the top-1 retrieved text. M2TM2T (motion-to-text) and M2MM2M (motion-to-motion) measure alignment through cosine similarity between embeddings of generated motions and ground truth texts or motions. In the speech-to-gesture category, FID−AFID-A similarly measures the realism of motion generated based on speech audio. Beat Consistency (BC) evaluates how well gestures synchronize with the rhythm and beats of the speech, while L1 Diversity (L1Div) quantifies gesture diversity by calculating the average L1 distance between multiple gesture clips. This comprehensive set of metrics thoroughly evaluates our method across key dimensions.

Comparison with Baselines

Baseline Setting. To effectively evaluate the performance of our method, we propose a series of baselines inspired by reasonable settings and prior works. These include: Weighted Sum, a method that follows [70] by combining the predictions of text- and audio-conditioned models through weighted sum; Pseudo-Text, a method that follows [37] by using pseudo text descriptions of a speaker’s speech as the text condition during training; and SynTalker [12], an open-source multi-modal motion generation method performing weighted sum in body-part-specific latent space. Notably, because these baselines cannot handle asynchronous conditions, we align the text time intervals with the audio time intervals to construct synchronous conditions for all baseline methods. The results are reported in Table 2.

Text2MotionSpeech2Gesture
FID+FID+ ↓R1R1 ↑M2TM2T ↑M2MM2M ↑FID−AFID-A ↓BCBC ↑L1divL1div ↑
GT (Ground Truth)0.00040.00.7811.000---
Weighted Sum [70]1.3356.80.5460.5372.172.204.08
Pseudo-Text [37]1.5932.20.5110.503<u>2.22</u>2.556.43
SynTalker [12]<u>0.985</u><u>9.8</u><u>0.601</u><u>0.603</u>6.602.959.12
MOCO0.86224.60.6490.6393.83<u>2.72</u><u>8.62</u>

Table 2. Comparison with baselines. We highlight the best performance in bold and use underlining to indicate the second-best performance.

ConditionStart TimeEnd TimeBody Part
a man walks0.08.0legs, spine
wave hands7.011.0left arm, right arm
speak:1_wayne_0_103_103,$9248$998082.5788.238left arm, right arm, head
speak:1_wayne_0_103_103,$107552$1960649.72215.254left arm, right arm, head

Table 2.

Result Analysis. As shown in the table, our proposed method, MOCO, exhibits robust performance across both sets of metrics, delivering competitive results in both text-to-motion and speech-to-gesture tasks simultaneously. This underscores the effectiveness of MOCO in generating condition-aligned motions when multi-modal conditions are provided concurrently.

Weighted Sum and Pseudo-Text achieves good results on speech-to-gesture metrics but poor performance on text-to-motion metrics, indicating their limited ability to handle multi-modal data. We explain this further in Appendix E.

SynTalker also notices the importance of separating different body parts, resulting in an improvement in text-to-motion metrics compared to Weighted Sum and Pseudo-Text. However, its insufficient decoupling of conditions and suboptimal joint latent space—negatively impacted by repetitive patterns in the speech-to-gesture data—results in limited expressiveness and unstable movements. This is reflected in its less competitive performance in R1 and FID-A.

Beyond the baselines above, STMC [49] is particularly relevant, as its per-step composition paradigm inspired our multi-modal decoupled denoising. Since STMC is text-only and does not natively accept speech audio, we adapt it to our concurrent setting and conduct a dedicated comparison, including a user study (Appendix G.1 and Appendix G.5). There, STMC attains competitive text-to-motion scores, yet MOCO produces more natural and better-synchronized motion, as confirmed by quantitative metrics, additional naturalness metrics, and the subjective user study.

Ablation Study

To assess the impact of key designs within MOCO, we conduct an ablation study presented in Table 3. This study systematically examines the effects of body masking (Body Mask), weight sharing (Share Weight), and combination at each step (Per-Step Comb.) on the model’s performance across text-to-motion and speech-to-gesture metrics. In addition, we introduce the Transition Smoothness Ratio (TSR) to measure temporal coherence at segment boundaries, defined as the ratio between the mean frame-to-frame speed within transition windows and the mean speed in non-transition regions.

MethodShare WeightBody Mask 1−I1-IPer-Step Comb.Text2Motion FID+ ↓\downarrowText2Motion R1 ↑\uparrowText2Motion M2T ↑\uparrowSpeech2Gesture FID-A ↓\downarrowSpeech2Gesture BC ↑\uparrowSpeech2Gesture L1div ↑\uparrowTSR ↓\downarrow
GT---0.00040.00.781----
MOCO×\timeshead, arms✓\checkmark0.86224.60.6493.832.728.620.90
Variant 1×\timeshead, arms, spine✓\checkmark0.92122.40.6343.862.818.350.90
Variant 2×\timesspine, legs✓\checkmark1.2347.90.5542.752.195.110.95
Variant 3✓\checkmarkhead, arms✓\checkmark0.86622.10.6564.242.688.950.85
Variant 4×\timeshead, arms×\times0.83224.40.6453.992.278.601.17

Table 3. Ablation study on key designs within MOCO. We highlight the best performance in bold and use underlining to indicate the second-best performance achieved by our method.

Body Masking. In Variants 1 and 2, we test our hypothesis that speech audio guides upper-body motion (head and arms) while text descriptions influence lower-body movements (spine and legs). In Variant 1, we expand the body mask to include the spine along with the head and arms (Body Mask = head, arms, spine). This modification results in a worse FID+ and a slight decrease in R1, indicating a decline in text-to-motion performance. Moreover, it does not produce significant improvements in speech-to-gesture metrics, suggesting that including the spine in the body mask fails to enhance gesture generation and instead compromises text-driven motion performance.

Variant 2 further adjusts the body mask to include the legs and spine (Body Mask = legs, spine), which significantly degrades text-to-motion metrics and Beat Consistency. This mainly stems from a mismatch: the texts in our multimodal benchmark describe diverse actions such as “walk,” “sit,” and “turn right,” whereas the speech-to-gesture data are mostly standing gestures. Controlling lower-body motion with audio makes it hard to align with these texts, while controlling upper-body motion with text complicates beat alignment. Although Variant 2 yields a notable improvement in FID-A—reflecting a bias in the speech-to-gesture data toward standing-in-place motions—its overall performance still degrades.

In contrast, our original method (Body Mask == head, arms) effectively balances the influences of both text and audio inputs. By assigning the upper body to be guided by audio and the lower body by text, we achieve superior results across both text-to-motion and speech-to-gesture metrics. This demonstrates the advantage of our approach in producing coherent and contextually appropriate motions that align well with the provided conditions.

Weight Sharing. In Variant 3, we enable weight sharing (Share Weight == ✓\checkmark), following previous multi-modal methods [12, 37, 70], while keeping the body mask and transition method unchanged. Compared to the full MOCO model (without weight sharing), enabling weight sharing results in worse performance across several metrics, including R1 and FID-A. This decline suggests that sharing weights between modalities may limit the model’s ability to capture modality-specific nuances, thereby reducing its effectiveness in generating accurate and realistic motions for both text-to-motion and speech-to-gesture tasks.

Combination at Each Step. Variant 4 performs body-part combination only at the final denoising step t=0t = 0 (Per-Step Comp. == ×\times). Lacking iterative coordination to aggregate holistic motion information, this baseline exhibits temporal discontinuities (reflected by high TSR) and reduced body coherence. A user study (Sec. 4.6) and qualitative results in the supplementary videos further confirm that MOCO produces more natural and synchronized motions.

Qualitative Analysis

To clearly illustrate the overall performance of MOCO, we visualize four generated samples and their corresponding conditions in Figure 4. The lighter color of the mesh and the text description background indicates the start of each sequence, while the darker color indicates the end. These results exhibit natural speech-driven upper-body gestures synchronized with diverse lower-body motions such as jogging, walking, and sitting, demonstrating MOCO’s ability to generate coherent and realistic motions that closely align with the given multimodal signals. A detailed analysis of these examples is provided in the figure caption. Additional qualitative results, including comparisons with baselines, are included in the supplementary materials.

Qualitative results showing four samples generated by MOCO

Figure 4. Qualitative results. We visualize four samples generated by MOCO. Darker colors represent later points in time. The results demonstrate that MOCO is capable of generating coherent and realistic motions that highly align with the given multi-modal control signals. Figures (a), (b), and (d) present natural speech gestures coordinated with various lower-body movements as specified by the text inputs, such as jogging, walking in a circle, turning right, running, and so on. Figure (c) displays natural speech movements while sitting down. Figure (d) reveals a limitation of MOCO. When standing up or sitting down, the foot should remain stationary. However, the foot highlighted in the red box slides, leading to unrealistic results.

User Study

Table 4 presents the results of a user study comparing MOCO with three methods. Specifically, we evaluate MOCO against Pseudo-Text and Weighted Sum, introduced in Sec. 4.3, and Combine Once (Variant 4 in Sec. 4.4), to assess overall performance in text following and audio beat synchronization. Additional experimental details are provided in Appendix G.4.

BaselineBetter Text Following (%)Better Beat Synchronization (%)
NeitherBaselineMOCO (Ours)NeitherBaselineMOCO (Ours)
Pseudo-Text1.00.099.013.012.574.5
Weighted Sum0.00.0100.013.53.583.0
BaselineBetter Body Coherence (%)Better Temporal Fluidity (%)
NeitherBaselineMOCO (Ours)NeitherBaselineMOCO (Ours)
Combine Once (Variant 4)12.017.071.012.642.644.8

Table 4. User study results comparing MOCO against baseline methods.

As shown in the table, MOCO achieves significant advantages over both Pseudo-Text and Weighted Sum, demonstrating the effectiveness of our decoupled denoising strategy in generating motion aligned with multi-modal conditions. Furthermore, when compared to Combine Once, our method was rated significantly higher in both body coherence and temporal fluidity. This indicates that MOCO does more than merely combine different body parts controlled by separate conditions; it ensures that each body part aligns with its corresponding condition while enhancing coordination among all body parts. To further strengthen these findings, we additionally conduct a larger-scale pairwise preference study with 51 participants comparing MOCO against STMC, where MOCO is preferred with statistical significance on audio synchronization and naturalness (p<0.05p < 0.05); full protocol and results are reported in the supplementary material.

Discussion and Limitations

MOCO achieves effective multi-modal motion generation, but its fixed body-part assignment can be suboptimal when control signals conflict. For instance, when text descriptions specify arm movements (e.g., “wave hands”), this rigid separation leaves the upper body solely governed by audio, preventing the model from following text instructions involving arm motions. To mitigate this limitation, MOCO allows users to configure body-part control through the motion timeline (see Appendix D), enabling text to drive upper-body motion when needed.

Our multi-modal benchmark is also deliberately focused on compositional consistency under controlled settings rather than open-ended semantic coverage, enabling precise and repeatable evaluation of cross-modal integration but not claiming generalization to arbitrary text–audio pairs. Future work may explore adaptive body-part assignment or residual blending that dynamically handles spatially overlapping actions, moving beyond fixed spatial rules toward context-aware fusion while keeping the interpretability of decoupled generation.

Conclusion

This study presents MOCO, a novel diffusion-based framework to generate realistic and coherent holistic body motions from multi-modal control signals, including text descriptions, speech audio, and trajectory data. Our key innovation lies in a decoupled denoising process where, during each denoising step, the model independently generates motions for each modality and assembles them according to predefined spatial rules. This approach ensures that the generated motion is closely aligned with each condition while producing realistic and coherent whole-body movements. Experimental results demonstrate that our approach delivers state-of-the-art performance both qualitatively and quantitatively, advancing the field of multi-modal controlled motion generation.

Acknowledgments. This work was supported by the National Natural Science Foundation of China under Grants 62476099 and 62076101, Guangdong Basic and Applied Basic Research Foundation under Grants 2024B1515020082 and 2023A1515010007, and the TCL Young Scholars Program.

References

Body-Part Motion Statistics

A core design principle of MOCO is the spatial decomposition of multi-modal conditioning: text descriptions govern lower-body motion, while speech audio drives upper-body dynamics. This appendix provides a quantitative characterization of the motion distributions associated with each conditioning modality within the scope of our target task, offering an empirical basis for understanding why this decomposition is well-suited to the concurrent text-and-speech control scenario.

Task-Scoped Text Conditions and Motivation for Generated Samples

Our target application involves concurrent control by both speech audio and text descriptions. In this setting, text conditions are expected to specify body-level actions that are spatially complementary to co-speech gestures—that is, actions primarily activating the lower body without conflicting with concurrent upper-body gesture generation. We therefore focus our analysis on a curated set of 40 text descriptions centered on locomotion and postural transitions (see Appendix H), which represent the intended operating domain of the text conditioning signal in MOCO. These descriptions are representative of the text conditions under which MOCO is designed to operate, and do not aim to characterize the full diversity of text-driven motion.

To obtain the motion distribution for these curated descriptions, we generate 1,000 sequences using a pretrained text-to-motion model (MDM [59]) trained on HumanML3D [20], as these specific lower-body–oriented descriptions do not have corresponding ground-truth annotations in existing datasets. This approach allows us to directly estimate the body-part activation patterns associated with the text conditions employed in our multi-modal benchmark.

Setup

Both data sources are represented using the unified 205-dimensional SMPL-X body parameterization. Body parts are defined via fixed joint indices, grouped into five regions: left arm, right arm, head, spine, and legs. For each motion sequence, we compute the temporal variance of each body-part region as a proxy for motion activity, and normalize across all five parts so that they sum to 100% per sequence. We further aggregate upper-body parts (head, left arm, right arm) and lower-body parts (spine, legs), and define the lower-body ratio as:

rlower=VarlowerVarupper+Varlower,(11)r_{\text{lower}} = \frac{\mathrm{Var}_{\text{lower}}}{\mathrm{Var}_{\text{upper}} + \mathrm{Var}_{\text{lower}}}, \tag*{(11)}

where Varupper\mathrm{Var}_{\text{upper}} and Varlower\mathrm{Var}_{\text{lower}} denote the aggregated temporal variance of upper- and lower-body joints, respectively. This normalization eliminates confounds from sequence length and global motion scale, enabling a fair comparison across datasets.

Statistics are computed over two sources:

  • Generated (text-driven): 1,000 sequences produced by a pretrained text-to-motion diffusion model conditioned on our curated set of 40 locomotion- and posture-oriented descriptions.

  • BEAT2 (speech-driven): 5,000 sequences sampled from the BEAT2 training set [40].

Results

Quantitative results are reported in Table 5. Under our task-scoped text conditions, the generated motions exhibit a lower-body variance share of 62.0%, indicating that locomotion- and posture-oriented descriptions predominantly activate the legs and spine. In contrast, BEAT2 speech-gesture sequences are strongly upper-body dominant, with 85.1% of motion variance concentrated in the head and arms—consistent with the expressive, gesture-rich nature of co-speech motion. These complementary distributions provide empirical support for our proposed spatial decomposition.

SourceUpper MeanVarLower MeanVarLower / UpperLower-Body Ratio
Generated (text-driven)0.003870.006631.71×62.0%
BEAT2 (speech-driven)0.005870.001640.28×14.9%

Table 5. Body-part motion statistics for text-driven generated sequences and BEAT2 speech-gesture data. Temporal variance is computed per body-part region and normalized within each sequence.

Theoretical Analysis for Decoupled Denoising

Our proposed MOCO relies on the assumption that the joint conditional probability p(bt−1∣ctext,caudio,bt)p(b_{t-1} \mid c_{\text{text}}, c_{\text{audio}}, b_t) can be approximated by p(bt−1,lower∣ctext,bt)⋅p(bt−1,upper∣caudio,bt)p(b_{t-1,\text{lower}} \mid c_{\text{text}}, b_t) \cdot p(b_{t-1,\text{upper}} \mid c_{\text{audio}}, b_t), expressed as:

p(bt−1∣ctext,caudio,bt)≈p(bt−1,lower∣ctext,bt)⋅p(bt−1,upper∣caudio,bt)(12)\begin{aligned} p(b_{t-1} \mid c_{\text{text}}, c_{\text{audio}}, b_t) &\approx p(b_{t-1,\text{lower}} \mid c_{\text{text}}, b_t) \cdot p(b_{t-1,\text{upper}} \mid c_{\text{audio}}, b_t) \tag*{(12)} \end{aligned}

where btb_t denotes the motion at denoising step tt, composed of upper-body motion bt,upperb_{t,\text{upper}} and lower-body motion bt,lowerb_{t,\text{lower}}.

p(bt−1∣ctext,caudio,bt)=p(bt−1,lower,bt−1,upper∣call), where call={ctext,caudio,bt}=p(bt−1,lower∣call)⋅p(bt−1,upper∣call,bt−1,lower)(13)p(b_{t-1} \mid c_{\text{text}}, c_{\text{audio}}, b_t) = p(b_{t-1,\text{lower}}, b_{t-1,\text{upper}} \mid c_{\text{all}}), \text{ where } c_{\text{all}} = \{c_{\text{text}}, c_{\text{audio}}, b_t\} = p(b_{t-1,\text{lower}} \mid c_{\text{all}}) \cdot p(b_{t-1,\text{upper}} \mid c_{\text{all}}, b_{t-1,\text{lower}}) \tag*{(13)}
≈p(bt−1,lower∣call)⋅p(bt−1,upper∣call)(14)\approx p(b_{t-1,\text{lower}} \mid c_{\text{all}}) \cdot p(b_{t-1,\text{upper}} \mid c_{\text{all}}) \tag*{(14)}
≈p(bt−1,lower∣call∖{caudio})⋅p(bt−1,upper∣call∖{ctext})=p(bt−1,lower∣ctext,bt)⋅p(bt−1,upper∣caudio,bt).(15)\begin{aligned} \approx p(b_{t-1,\text{lower}} \mid c_{\text{all}} \setminus\{c_{\text{audio}}\}) \cdot p(b_{t-1,\text{upper}} \mid c_{\text{all}} \setminus\{c_{\text{text}}\}) \\ &= p(b_{t-1,\text{lower}} \mid c_{\text{text}}, b_t) \cdot p(b_{t-1,\text{upper}} \mid c_{\text{audio}}, b_t). \tag*{(15)} \end{aligned}

The first approximation occurs in the transition from Eq. (13) to Eq. (14). Here, we approximate:

p(bt−1,upper∣call,bt−1,lower)=p(bt−1,upper∣ctext,caudio,bt,upper,bt,lower,bt−1,lower)≈p(bt−1,upper∣ctext,caudio,bt,upper,bt,lower)=p(bt−1,upper∣call).\begin{aligned} p(b_{t-1,\text{upper}} \mid c_{\text{all}}, b_{t-1,\text{lower}}) &= p(b_{t-1,\text{upper}} \mid c_{\text{text}}, c_{\text{audio}}, b_{t,\text{upper}}, b_{t,\text{lower}}, b_{t-1,\text{lower}}) \\ &\approx p(b_{t-1,\text{upper}} \mid c_{\text{text}}, c_{\text{audio}}, b_{t,\text{upper}}, b_{t,\text{lower}}) \\ &= p(b_{t-1,\text{upper}} \mid c_{\text{all}}). \end{aligned}

This approximation assumes that btb_t already encapsulates sufficient information about bt−1b_{t-1}, allowing us to neglect the influence of bt−1,lowerb_{t-1,\text{lower}} when estimating bt−1,upperb_{t-1,\text{upper}}. This simplification is justified by the proximity of the diffusion steps and the strong correlation between the states at steps tt and t−1t-1.

The second approximation occurs in the transition from Eq. (14) to Eq. (15), where we decouple modality-specific influences:

p(bt−1,lower∣call)≈p(bt−1,lower∣call∖{caudio}),p(bt−1,upper∣call)≈p(bt−1,upper∣call∖{ctext}).\begin{aligned} p(b_{t-1,\text{lower}} \mid c_{\text{all}}) &\approx p(b_{t-1,\text{lower}} \mid c_{\text{all}} \setminus\{c_{\text{audio}}\}), \\ p(b_{t-1,\text{upper}} \mid c_{\text{all}}) &\approx p(b_{t-1,\text{upper}} \mid c_{\text{all}} \setminus\{c_{\text{text}}\}). \end{aligned}

This approximation leverages the observation that text input (ctextc_{\text{text}}) primarily influences lower-body movements (e.g., walking or shifting stance), while audio input (caudioc_{\text{audio}}) predominantly affects upper-body movements (e.g., gestures or facial expressions). By excluding caudioc_{\text{audio}} from the conditioning set for bt−1,lowerb_{t-1,\text{lower}} and ctextc_{\text{text}} for bt−1,upperb_{t-1,\text{upper}}, we ensure the conditioning focuses on the most relevant modality for each body part.

Leveraging LLM for Motion Planning

In summary, our prompt instructing the LLM to generate a motion timeline consists of several key components:

  1. Purpose: Clearly state the objective of generating a motion timeline based on text and audio conditions.

  2. Output Formats: Specify the formats that the LLM should adhere to when producing output.

  3. Guidelines for Generating Timelines: Outline several essential rules that the LLM must follow.

  1. Examples of Timelines: Provide a pair of example timelines, showcasing both an effective and a less effective version.

Additionally, to enhance the precision and completeness of the LLM’s reasoning, we append the phrase “please reason step by step” at the end of the dialogue.

To illustrate this process more clearly, we provide a complete dialogue record with GPT-4o mini in motion-planning-dialogue.pdf. In this dialogue, the input conditions are processed by the LLM as follows:

Explanation:

  • Text Description: “sit down on a chair”

    • Start and End Time in Overall Motion: # 0.0 # 5.0

    • Controlled Body Parts: # legs # spine

  • Audio Input File: speak:5_stewart_0_1_1

    • Start and End Frames in Audio File: $785440$911328

    • Start and End Time in Overall Motion: # 5.0 # 12.85

    • Controlled Body Parts: # head # arms

Notable, the LLM may occasionally produce calculation errors, such as miscalculating the duration of the second audio segment. However, these issues can be easily resolved through code refinement. The motion generated by our MOCO based on the refined timeline is provided in motion-planning-result.mp4.

Conflict Resolution and Flexibility in Multi-Modal Control

MOCO ’s default configuration assigns upper-body control to audio signals and lower-body control to text descriptions. However, there are scenarios where users may wish to explicitly control specific upper-body motions, such as waving hands while talking. To clarify how our method handles this case, we provide the following example:

In this example, a conflict arises between the upper-body motion cues from the text description “waves hands” and the speech audio during the overlapping period from 2.578 seconds to 8.0 seconds. To resolve such conflicts, our method relies on a predefined timeline that gives precedence to the speech audio in governing upper-body motion during this interval. Consequently, the ‘waves hands” motion is suppressed between 2.578 and 8.0 seconds.

While this illustrates a limitation of our default setup, the flexibility of MOCO offers a way to address such conflicts by enabling finer-grained user control. If users wish to control upper-body motions using text during speech, we can simply adjust the timeline as follows:

After modifying the timeline, the text command “wave hands” controls the arms from 7.0 to 11.0 seconds. Since the speech audio spans from 2.578 to 15.254 seconds, there is an overlap between 7.0 and 11.0 seconds during which both audio and text attempt to control the arms. According to our conflict resolution rules, the condition with fewer controlled body parts takes precedence. In this case, the text command “wave hands” governs the arms during the overlapping period.

This example demonstrates the flexibility of our method. Although our default configuration prioritizes audio for the upper body and text for the lower body, users can easily reconfigure these priorities through timeline adjustments to achieve their desired behavior. The corresponding visualization is provided in the supplementary materials as flexibility.mp4.

Limitations of Weighted Sum in Multi-Modal Motion Generation

To understand why the Weighted Sum method achieves good results in speech-to-gesture metrics but poor performance in text-to-motion metrics, we conducted the following experiments.

Given both speech and text inputs, we used only the text-to-motion model to update the motion. At each denoising step tt, we computed the difference difft\mathrm{diff}_t between the speech-to-gesture model’s prediction—based on the speech input and the current motion from the text-to-motion model—and the current motion from the text-to-motion model. This difference quantifies how much the speech-to-gesture model perceives a mismatch between the speech condition and the current motion. Conversely, when we used only the speech-to-gesture model to update the motion, the calculated difference indicated how much the text-to-motion model perceived a mismatch between the text condition and the current motion. A larger difference suggests a greater mismatch and that the model will update the motion more aggressively.

We recorded these differences in both scenarios and divided them into whole body, arms, and legs for clearer illustration. Comparing Figure 5 (a) and (b), as well as Figure 5 (c) and (d), we found that the differences calculated by the speech-to-gesture model are larger than those by the text-to-motion model. This indicates that the speech-to-gesture model adjusts the motion more aggressively based on its conditions than the text-to-motion model does. This explains why, when using the weighted sum method, the generated result closely resembles that produced entirely by the speech-to-gesture model.

Comparison of differences calculated by the speech-to-gesture and text-to-motion models during motion updates

Figure 5. Comparison of differences calculated by the speech-to-gesture and text-to-motion models during motion updates. “T” denotes using a text-to-motion model to update motion, while “A” denotes using a speech-to-gesture model to update motion. The results show that the speech-to-gesture model computes larger differences than the text-to-motion model, indicating it adjusts the motion more aggressively based on the conditions. This explains why the weighted sum method attains strong performance on speech-to-gesture metrics but poor performance on text-to-motion metrics. Additionally, when the text condition is “sitting,” the speech-to-gesture model calculates larger differences in the legs than in the arms. This is counterintuitive and could be attributed to data bias in the speech-to-motion dataset. Best viewed in color.

Furthermore, by comparing Figure 5 (a) and (c), which have different text conditions, we observe that when the text condition is “sitting,” the differences calculated by the speech-to-gesture model in the legs are larger than in the arms. This is counterintuitive since speech is typically associated with upper-body gestures rather than lower-body movements. Conversely, when the text condition is “standing,” the differences in the legs are smaller than in the arms, aligning with expectations. This phenomenon may be attributed to data bias in the speech-to-motion dataset, where most motions are performed in standing positions.

These observations reveal the limitations of weighted sum in multi-modal motion generation and suggest the validity of our proposed decoupled denoising process.

Computational Complexity

Our framework, MOCO, comprises four transformer-based denoisers: GT2MG_{\mathrm{T2M}} for text-to-motion (T2M), GS2GG_{\mathrm{S2G}} for speech-to-gesture (S2G), GT2VG_{\mathrm{T2V}} for trajectory-to-velocity (T2V), and GS2DG_{\mathrm{S2D}} for speech-to-details (S2D), which manages facial expressions and hand poses. To clearly illustrate the computational complexity of MOCO, we present various metrics, including the number of parameters, model size, FLOPs, and inference time on a single NVIDIA 4090 GPU, as shown in Table 6.

Parameters (M)Model Size (MB)FLOPs (G)Inference Time (ms/frame)
GT2MG_{\mathrm{T2M}}27.01103.025.192.26
GS2GG_{\mathrm{S2G}}36.86140.626.724.30
GT2VG_{\mathrm{T2V}}0.341.310.066.20
GS2DG_{\mathrm{S2D}}36.94140.946.744.37

Table 6. Complexity of each denoiser of MOCO.

As indicated in the table, our framework is overall lightweight and sufficiently fast. Specifically, the speech-to-gesture denoiser GS2GG_{\mathrm{S2G}} and the speech-to-details denoiser GS2DG_{\mathrm{S2D}} are relatively larger than the other denoisers due to additional cross-attention parameters. In contrast, the trajectory-to-velocity denoiser GT2VG_{\mathrm{T2V}} is the most lightweight module, featuring fewer hidden state dimensions and transformer layers because the task it handles involves low-dimensional data. However, the introduction of a guidance mechanism for more accurate predictions results in GT2VG_{\mathrm{T2V}} having the longest inference time.

Finally, to generate the motion sequences for a 35-second demo video consisting of nine clips under different conditions and with a total duration of 54 seconds, our method completed the body motion generation task in only 3.72 seconds. This fast generation time highlights the potential of our approach for real-time applications.

For a like-for-like comparison against competing methods, we additionally report end-to-end generation speed (seconds per sequence) under the concurrent multi-modal setting in Table 7 (Appendix G.1): MOCO is the fastest at 0.69 s/seq, ahead of STMC (0.75 s/seq) and the StableMoFusion-backbone variant (0.86 s/seq), confirming that MOCO attains its naturalness and synchronization advantages without incurring additional inference cost.

MethodNaturalness MC↑Naturalness Jitterspine ↓Text2Motion FID+↓Text2Motion R1↑Speech2Gesture FID-A↓Speech2Gesture BC↑Speed s/seq↓
GT−1.950.590.00040.0–––
MOCO (Ours)−3.390.600.86224.63.832.720.69
STMC [49]−3.840.630.68427.54.442.220.75
StableMoFusion [30]−3.910.690.75227.24.832.920.86

Table 7. Comparison with per-step composition (STMC) and a stronger diffusion backbone (StableMoFusion). Best among generative methods is in bold.

Additional Experiments

Comparison with STMC and Stronger Backbones

This section presents experiments comparing MOCO against two additional baselines, with results reported in Table 7. We include STMC [49] because its singlemodality-per-step composition inspired our multi-modal decoupled denoising, and StableMoFusion [30] to examine whether a backbone that is stronger on text-to-motion transfers into improved motion quality in our concurrent control setting. To adapt STMC to the multi-modal setting while faithfully preserving its original design, we use self-attention to process both text and audio conditions, share parameters across denoisers, and rely on an LLM to automatically annotate the assignment of body parts. In addition to the text-to-motion (FID+, R1) and speech-to-gesture (FID-A, BC) metrics used in the main paper, we further report two naturalness-oriented metrics: MotionCritic (MC) [61], a learned motion-quality score aligned with human perception, and Jitterspine, the jerk of the spine joints at the boundary between the audio-controlled upper body and the text-controlled lower body, which quantifies how smoothly the two independently denoised regions are stitched together.

CFGOCL-BFGSLocation Average DifferenceLocation Goal DifferenceOrientation Average DifferenceOrientation Goal Difference
×\times×\times×\times0.561.210.711.26
✓\checkmark×\times×\times0.571.320.831.51
✓\checkmark×\times✓\checkmark0.110.200.711.28
×\times✓\checkmark×\times0.440.790.530.97
×\times×\times✓\checkmark0.080.130.591.10
×\times✓\checkmark✓\checkmark0.140.200.480.87

Table 8. Evaluation of trajectory control under classifier-free guidance (CFG), Omni-Control (OC) [66], and post-hoc L-BFGS optimization, on location (m) and orientation (rad). OmniControl is a dedicated spatial-control method used to replace our trajectory-to-velocity module GT2VG_{\mathrm{T2V}}; the first three rows therefore use GT2VG_{\mathrm{T2V}}.

We report the results transparently. On pure text-to-motion fidelity, STMC is in fact slightly stronger than MOCO, and StableMoFusion is likewise competitive on these metrics. However, MOCO achieves better speech-to-gesture fidelity, boundary smoothness, and MotionCritic score. These results suggest that, rather than being uniformly best on every isolated single-modality metric, MOCO delivers the most natural and best-synchronized motion under joint text-and-speech control, as further corroborated by the pairwise user study in Appendix G.5. We note that MC should be interpreted only as an auxiliary indicator, because it is trained predominantly on text-driven motion, whereas MOCO produces dense, expressive co-speech upper-body gestures that may lie partly outside its training distribution.

Evaluation of Trajectory Control

Tab. 8 evaluates trajectory control methodologies by assessing the effects of classifier-free guidance (CFG), OmniControl (OC) [66], and L-BFGS optimization on both location (meters) and orientation (radians). Here, OmniControl is a dedicated spatial-control method that replaces our lightweight trajectory-to-velocity module GT2VG_{\mathrm{T2V}}, whereas the other configurations build on GT2VG_{\mathrm{T2V}}. For each category, two metrics are reported: Average Difference, the mean deviation between the generated trajectory and the ground truth (GT), and Goal Difference, the discrepancy at the final point relative to the GT.

The post-hoc L-BFGS optimization is the single most effective component: applied on top of GT2VG_{\mathrm{T2V}}, it drastically reduces the location error while also modestly improving orientation. In contrast, CFG does not help and can even interfere with the optimization, slightly worsening both location and orientation. Replacing GT2VG_{\mathrm{T2V}} with OmniControl improves orientation, and OC+L-BFGS achieves the overall best orientation—but at the cost of a larger location error and additional inference time. We therefore retain GT2VG_{\mathrm{T2V}}+L-BFGS as our default for its better location accuracy and efficiency, and note that adopting a dedicated spatial controller is a worthwhile direction when orientation fidelity is prioritized.

We further emphasize, in the interest of transparency, that the absolute orientation accuracy of MOCO is still not fully satisfactory. We attribute this to a deliberate design trade-off: MOCO adopts a motion representation chosen to facilitate temporal stitching, which is prone to accumulated error along the trajectory—hence the orientation error remains comparatively large even after optimization.

Single Modality Performance

To demonstrate MOCO’s performance in single-modality scenarios, we trained it from scratch on HumanML3D for text-to-motion and on BEAT2 for speech-to-gesture, respectively, ensuring a fair comparison. The results, presented in Table 9 and Table 10, show that in the HumanML3D text-to-motion benchmark (Table 9), our model achieves performance comparable to the widely-used MDM. This outcome is expected since our text-to-motion denoiser, GT2MG_{\mathrm{T2M}}, is based on MDM. In the BEAT2 speech-to-gesture benchmark (Table 10), MOCO attains competitive performance compared to state-of-the-art methods.

MethodR-Precision Top 1R-Precision Top 2R-Precision Top 3FID↓MM Dist↓Diversity↑MM↑
Ground Truth0.5110.7030.7970.0022.9749.503-
T2M-GPT [74]0.4910.6800.7750.1163.1189.7611.856
MDM [59]--0.6110.5445.5669.5592.799
FineMoGen [77]0.5040.6900.7840.1512.9989.2632.696
MoMask [19]0.5210.7130.8070.0452.958-1.241
LMM-Tiny [76]0.4960.6850.7850.4153.0879.1761.465
LMM-Large [76]0.5250.7190.8110.0402.9439.8142.683
SynTalker [12]0.3750.5640.6814.3854.4999.374-
MOCO (Ours)0.4340.6180.7200.5303.5639.8562.663

Table 9. Quantitative results of text-to-motion generation on the HumanML3D test set.

MethodFGD↓BCDiversity↑MSE↓LVD↓
FaceFormer [17]---7.7877.593
CodeTalker [67]---8.0267.766
S2G [18]28.154.6835.971--
Trimodal [73]12.415.9337.724--
HA2G [43]12.326.7798.626--
DisCo [39]9.4176.4399.912--
CaMN [41]6.6446.76910.86--
DiffStyleGesture [69]8.8117.24111.49--
TalkShow [71]6.2096.94713.477.7917.771
EMAGE [40]5.5127.72413.067.6807.556
ProbTalk [44]6.1708.09910.438.9908.385
SynTalker [12]6.4137.97112.72--
MOCO (Ours)5.5437.08914.057.2857.573

Table 10. Quantitative results of speech-to-gesture generation on the BEATX test set.

Details of User Study

To construct the questionnaire for the user study, we first generated 20 videos using each method, including MOCO (Ours), Pseudo-Text, Weighted Sum, and Combine Once. The videos generated by MOCO were then concatenated with the other videos, resulting in 60 pairs of comparison videos. The user study involved 10 participants, each of whom evaluated all 60 pairs of comparison videos. Participants were asked to select their preferred video or indicate that neither was better based on one of the following criteria: text following, beat synchronization, body coherence, and temporal fluidity. The first 40 video comparisons involved MOCO versus Pseudo-Text and MOCO versus Weighted Sum, focusing on text following and beat synchronization. The last 20 video comparisons involved MOCO versus Combine Once, focusing on body coherence and temporal fluidity. The statistical results are presented in the main paper.

Extended Pairwise-Preference Study against STMC.

The study above compares MOCO against the weighted-sum, pseudo-text, and combine-once baselines. To further corroborate the metric-level comparison with STMC reported in Appendix G.1 through human perception, we additionally conducted a larger-scale pairwise-preference study against STMC [49] (the same adapted STMC baseline as in Appendix G.1). This study involved 51 participants drawn from a bachelor’s-level (undergraduate) background, none of whom were involved in the project; each evaluated 60 comparison videos. We report participant background explicitly here for transparency, as the demographic composition of evaluators can influence perceptual judgments. For each video pair, participants chose the preferred clip—or “no preference”—along three criteria: text alignment, audio synchronization, and naturalness.

The results are summarized in Table 11. MOCO is preferred over STMC on all three criteria, with the margin reaching statistical significance for audio synchronization (47.1% vs. 39.4%, p=0.02p = 0.02) and naturalness (46.1% vs. 39.8%, p=0.03p = 0.03), using a two-sided binomial test with ties excluded. For text alignment the preference favors MOCO but does not reach significance (45.9% vs. 42.0%, p=0.11p = 0.11), which is consistent with the metric-level observation in Appendix G.1 that STMC is competitive on text-to-motion metrics. Taken together, these perceptual results reinforce our central claim: under concurrent text-and-speech control, MOCO produces motion that human viewers judge to be better synchronized and more natural, even where it does not dominate on isolated text-to-motion metrics.

CriterionPrefer MOCONo PreferencePrefer STMCpp-value
Text Alignment45.9%12.1%42.0%0.11
Audio Synchronization47.1%13.5%39.4%0.02*
Naturalness46.1%14.1%39.8%0.03*

Table 11. Extended pairwise-preference user study comparing MOCO against STMC [49] under concurrent text-and-speech control. Each cell reports the fraction of comparisons in which participants preferred MOCO, expressed no preference, or preferred STMC. The pp-value is from a two-sided binomial test on the preference between the two methods (ties excluded); * denotes statistical significance at p<0.05p < 0.05.

Details of Multi-Modal Benchmark

To effectively evaluate our proposed task, we developed a multi-modal benchmark comprising 1,000 test clips, following the methodology outlined in [49]. Each test clip is automatically generated and includes two text descriptions and two audio clips.

For the text descriptions, we manually curated a set of 40 texts focusing on lower-body movements commonly associated with speech delivery or conversation. These descriptions provide the necessary context for evaluating the corresponding movements within the benchmark. Regarding the audio clips, we selected recordings from the BEAT2 dataset, specifically choosing eight speakers with speaker IDs below 10. These audio files were segmented into clips using a Voice Activity Detector (VAD), resulting in 694 audio clips with an average duration of 9.14 seconds.

The 1,000 test clips were generated through an automated process that utilizes the curated text descriptions and audio clips. For each test clip, two text descriptions are randomly selected and assigned random durations. Subsequently, two neighboring audio clips are randomly chosen. The start times for both the text and audio intervals are determined randomly, allowing the sequence to commence with either text or audio. This process results in the creation of four intervals that correspond to the selected text descriptions and audio clips.

Optional text descriptions are listed here:

walk in a circle clockwise
walk in a circle counterclockwise
walk in a quarter circle to the left
walk in a quarter circle to the right
turn 180 degrees to the left on the left foot
turn 180 degrees to the left on the right foot
turn left
turn right
walk forwards
walk backwards
slowly walk forwards
slowly walk backwards
quickly walk forwards
quickly walk backwards
run
jogs forwards
jogs backwards
slowly walk in a circle
perform a squat
sit down
turn around then sit down in a chair
sit down then get back up and walk back
sit down for a moment
step back and sit down
sit down indian style
take a step to their right and sit down
sit criss cross
sit down on the ground and cross their legs
squat down
sit on a high object
sit on a barstool and rest their legs on the stool
take a large step and sits on a stool
  1. get down on their knees

  1. sit on the ground with his legs extended in front of him

  1. walk up to a backwards chair and sit down on it with legs outstretched

  1. sit down and adjust themselves

  1. sit down and swap their legs crossing back and forth

  1. sit and lie down on a lounge chair

  1. sit down and lean on the chair

  1. sits very still in the chair

References

  1. [1]Alexanderson, S., Nagy, R., Beskow, J., Henter, G.E.: Listen, denoise, action! audio-driven motion synthesis with diffusion models. ACM Transactions on Graphics 42(4), 1–20 (2023)DOI
  2. [2]Ao, T., Gao, Q., Lou, Y., Chen, B., Liu, L.: Rhythmic gesticulator: Rhythm-aware co-speech gesture synthesis with hierarchical neural embeddings. ACM Transactions on Graphics 41(6), 1–19 (2022)DOI
  3. [3]Ao, T., Zhang, Z., Liu, L.: Gesturediffclip: Gesture diffusion model with clip latents. ACM Transactions on Graphics 42(4), 1–18 (2023)arxiv.org/abs/2303.14613
  4. [4]Athanasiou, N., Petrovich, M., Black, M.J., Varol, G.: Teach: Temporal action composition for 3d humans. In: International Conference on 3D Vision (3DV). pp. 414–423. IEEE (2022)arxiv.org/abs/2209.04066
  5. [5]Athanasiou, N., Petrovich, M., Black, M.J., Varol, G.: Sinc: Spatial composition of 3d human motions for simultaneous action generation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 9984–9995 (2023)
  6. [6]Bae, J., Hwang, I., Lee, Y.Y., Guo, Z., Liu, J., Ben-Shabat, Y., Kim, Y.M., Kapadia, M.: Less is more: Improving motion diffusion models with sparse keyframes. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2025)arxiv.org/abs/2503.13859
  7. [7]Baevski, A., Zhou, Y., Mohamed, A., Auli, M.: wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems (NeurIPS) 33, 12449–12460 (2020)
  8. [8]Barquero, G., Escalera, S., Palmero, C.: Celfusion: Latent diffusion for behavior-driven human motion prediction. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 2317–2327 (2023)
  9. [9]Barquero, G., Escalera, S., Palmero, C.: Seamless human motion composition with blended positional encodings. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 457–469 (2024)arxiv.org/abs/2402.15509
  10. [10]Cen, Z., Pi, H., Peng, S., Shen, Z., Yang, M., Zhu, S., Bao, H., Zhou, X.: Generating human motion in 3d scenes from text descriptions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 1855–1866 (2024)
  11. [11]Cen, Z., Pi, H., Peng, S., Shuai, Q., Shen, Y., Bao, H., Zhou, X., Hu, R.: Ready-to-react: Online reaction policy for two-character interaction generation. In: International Conference on Learning Representations (ICLR) (2025)arxiv.org/abs/2502.20370
  12. [12]Chen, B., Li, Y., Ding, Y.X., Shao, T., Zhou, K.: Enabling synergistic full-body control in prompt-based co-speech motion generation. In: Proceedings of the ACM International Conference on Multimedia. pp. 6774–6783 (2024)arxiv.org/abs/2410.00464
  13. [13]Chen, C., Zhang, J., Lakshmikanth, S.K., Fang, Y., Shao, R., Wetzstein, G., Fei-Fei, L., Adeli, E.: The language of motion: Unifying verbal and non-verbal language of 3d human motion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2025)
  14. [14]Chen, X., Jiang, B., Liu, W., Huang, Z., Fu, B., Chen, T., Yu, G.: Executing your commands via motion diffusion in latent space. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 18000–18010 (2023)
  15. [15]Cong, P., Wang, Z., Dou, Z., Ren, Y., Yin, W., Cheng, K., Sun, Y., Long, X., Zhu, X., Ma, Y.: Laserhuman: language-guided scene-aware human motion generation in free environment. arXiv preprint arXiv:2403.13307 (2024)arxiv.org/abs/2403.13307
  16. [16]Dai, W., Chen, L.H., Wang, J., Liu, J., Dai, B., Tang, Y.: MotionIcm: Real-time controllable motion generation via latent consistency model. In: European Conference on Computer Vision (ECCV). pp. 390–408. Springer (2024)
  17. [17]Fan, Y., Lin, Z., Saito, J., Wang, W., Komura, T.: FaceFormer: Speech-driven 3D facial animation with transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 18770–18780 (2022)
  18. [18]Ginosar, S., Bar, A., Kohavi, G., Chan, C., Owens, A., Malik, J.: Learning individual styles of conversational gesture. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 3497–3506 (2019)arxiv.org/abs/1906.04160
  19. [19]Guo, C., Mu, Y., Javed, M.G., Wang, S., Cheng, L.: Momask: Generative masked modeling of 3d human motions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 1900–1910 (2024)
  20. [20]Guo, C., Zou, S., Zuo, X., Wang, S., Ji, W., Li, X., Cheng, L.: Generating diverse and natural 3d human motions from text. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 5152–5161 (2022)
  21. [21]Guo, C., Zuo, X., Wang, S., Zou, S., Sun, Q., Deng, A., Gong, M., Cheng, L.: Action2motion: Conditioned generation of 3d human motions. In: Proceedings of the ACM International Conference on Multimedia. pp. 2021–2029 (2020)
  22. [22]Gupta, P., Verma, S., Grama, A., Bera, A.: Unified multi-modal interactive & reactive 3d motion generation via rectified flow. In: International Conference on Learning Representations (ICLR) (2026)
  23. [23]Harithas, S., Sridhar, S.: Motionglot: A multi-embodied motion generation model. In: International Conference on Learning Representations (ICLR) (2025)
  24. [24]Hassan, M., Choutas, V., Tzionas, D., Black, M.J.: Resolving 3d human pose ambiguities with 3d scene constraints. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 2282–2292 (2019)
  25. [25]Ho, J., Salimans, T.: Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598 (2022)arxiv.org/abs/2207.12598
  26. [26]Hong, S., Kim, C., Yoon, S., Nam, J., Cha, S., Noh, J.: Salad: Skeleton-aware latent diffusion for text-driven motion generation and editing. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 7158–7168 (2025)
  27. [27]Hosseyni, S.R., Rahmani, A.A., Seyedmohammadi, S.J., Seyedin, S., Mohammadi, A.: Bad: Bidirectional auto-regressive diffusion for text-to-motion generation. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 1–5. IEEE (2025)
  28. [28]Hu, L., Ye, Y., Xia, S.: Hmvlm: Human motion-vision-language model via moe lora. arXiv preprint arXiv:2511.01463 (2025)arxiv.org/abs/2511.01463
  29. [29]Huang, S., Wang, Z., Li, P., Jia, B., Liu, T., Zhu, Y., Liang, W., Zhu, S.C.: Diffusion-based generation, optimization, and planning in 3d scenes. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 16750–16761 (2023)
  30. [30]Huang, Y., Yang, H., Luo, C., Wang, Y., Xu, S., Zhang, Z., Zhang, M., Peng, J.: Stablemofusion: Towards robust and efficient diffusion-based motion generation framework. In: Proceedings of the ACM International Conference on Multimedia. pp. 224–232 (2024)
  31. [31]Jiang, Y., Liao, Q., Lu, Z.: Smoothsync: Dual-stream diffusion transformers for jitter-robust beat-synchronized gesture generation from quantized audio. arXiv preprint arXiv:2601.04236 (2026)arxiv.org/abs/2601.04236
  32. [32]Karunratanakul, K., Preechakul, K., Suwajanakorn, S., Tang, S.: Guided motion diffusion for controllable human motion synthesis. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 2151–2162 (2023)
  33. [33]Lee, T., Baradel, F., Lucas, T., Lee, K.M., Rogez, G.: T2lm: Long-term 3d human motion generation from multiple sentences. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 1867–1876 (2024)
  34. [34]Li, R., Yang, S., Ross, D.A., Kanazawa, A.: Ai choreographer: Music conditioned 3d dance generation with aist++. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 13401–13412 (2021)
  35. [35]Li, Z., Yuan, W., He, Y., Qiu, L., Zhu, S., Gu, X., Shen, W., Dong, Y., Dong, Z., Yang, L.T.: Lamp: Language-motion pretraining for motion generation, retrieval, and captioning. In: International Conference on Learning Representations (ICLR) (2025)
  36. [36]Liang, H., Zhang, W., Li, W., Yu, J., Xu, L.: Intergen: Diffusion-based multi-human motion generation under complex interactions. International Journal of Computer Vision pp. 1–21 (2024)
  37. [37]Ling, Z., Han, B., Wong, Y., Kangkanghali, M., Geng, W.: Mcm: Multi-condition motion synthesis framework for multi-scenario. arXiv preprint arXiv:2309.03031 (2023)arxiv.org/abs/2309.03031
  38. [38]Liu, D.C., Nocedal, J.: On the limited memory bfgs method for large scale optimization. Mathematical programming 45(1), 503–528 (1989)
  39. [39]Liu, H., Iwamoto, N., Zhu, Z., Li, Z., Zhou, Y., Bozkurt, E., Zheng, B.: Disco: Disentangled implicit content and rhythm learning for diverse co-speech gestures synthesis. In: Proceedings of the ACM International Conference on Multimedia. pp. 3764–3773 (2022)
  40. [40]Liu, H., Zhu, Z., Becherini, G., Peng, Y., Su, M., Zhou, Y., Iwamoto, N., Zheng, B., Black, M.J.: Emage: Towards unified holistic co-speech gesture generation via masked audio gesture modeling. arXiv preprint arXiv:2401.00374 (2023)arxiv.org/abs/2401.00374
  41. [41]Liu, H., Zhu, Z., Iwamoto, N., Peng, Y., Li, Y., Zhou, Y., Zhou, W., Bozkurt, E., Zheng, B.: Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis. In: European Conference on Computer Vision (ECCV). pp. 612–630. Springer (2022)
  42. [42]Liu, P., Song, L., Huang, J., Liu, H., Xu, C.: GestureIsm: Latent shortcut based co-speech gesture generation with spatial-temporal modeling. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2025)
  43. [43]Liu, X., Wu, Q., Zhou, H., Xu, Y., Qian, R., Lin, X., Zhou, X., Wu, W., Dai, B., Zhou, B.: Learning hierarchical cross-modal association for co-speech gesture generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 10462–10472 (2022)
  44. [44]Liu, Y., Cao, Q., Wen, Y., Jiang, H., Ding, C.: Towards variable and coordinated holistic co-speech motion generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 1566–1576 (2024)
  45. [45]Liu, Y., Cao, Q., Yi, H., Jiang, H., Ding, C.: Multi-modal controlled coherent motion synthesis. OpenReview (2024), https://openreview.net/forum?id=i5Gxilzk0u
  46. [46]Liu, Y., Chen, C., Yi, L.: Interactive humanoid: One full-body motion reaction synthesis with social affordance canonicalization and forecasting. ArXiv abs/2312.08983 (2023), https://api.semanticscholar.org/CorpusID: 266209846arxiv.org/abs/2312.08983
  47. [47]Ma, S., Cao, Q., Zhang, J., Tao, D.: Contact-aware human motion generation from textual descriptions. arXiv preprint arXiv:2403.15709 (2024)arxiv.org/abs/2403.15709
  48. [48]Peng, Y., Song, J.T., Jung, S., Liu, R., Liu, H., Chu, X., Liu, R., Wu, E., Koike, H., Kitani, K.: Dyadit: A multi-modal diffusion transformer for socially favorable dyadic gesture generation. arXiv preprint arXiv:2602.23165 (2026)arxiv.org/abs/2602.23165
  49. [49]Petrovich, M., Litany, O., Iqbal, U., Black, M.J., Varol, G., Peng, X.B., Rempe, D.: Multi-track timeline control for text-driven 3d human motion generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 1911–1921 (2024)
  50. [50]Qi, X., Zhang, H., Wang, Y., Pan, J., Liu, C., Sun, M., Xue, W., Zhang, S., Han, S., Liu, Q., et al.: Cocosture: Towards coherent co-speech 3d gesture generation in the wild. Information Fusion p. 103613 (2025)
  51. [51]Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning (ICML). pp. 8748–8763. PMLR (2021)
  52. [52]Sampieri, A., Palma, A., Spinelli, I., Galasso, F.: Length-aware motion synthesis via latent diffusion. In: European Conference on Computer Vision (ECCV). pp. 107–124. Springer (2024)
  53. [53]Sha, X., Zhang, L., Mashita, T., Chiba, N., Uranishi, Y.: 3dgespolicy: Phoneme-aware holistic co-speech gesture generation based on action control. arXiv preprint arXiv:2601.18451 (2026)arxiv.org/abs/2601.18451
  54. [54]Shafir, Y., Tevet, G., Kapon, R., Bermano, A.H.: Human motion diffusion as a generative prior. arXiv preprint arXiv:2303.01418 (2023)arxiv.org/abs/2303.01418
  55. [55]Siyao, L., Yu, W., Gu, T., Lin, C., Wang, Q., Qian, C., Loy, C.C., Liu, Z.: Bailando: 3D dance generation by actor-critic GPT with choreographic memory. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 11050–11059 (2022)
  56. [56]Sun, J., Zhang, Q., Duan, Y., Jiang, X., Cheng, C., Xu, R.: Prompt, plan, perform: Llm-based humanoid control via quantized imitation learning. In: IEEE International Conference on Robotics and Automation (ICRA). pp. 16236–16242. IEEE (2024)
  57. [57]Sun, M., Xu, C., Jiang, X., Liu, Y., Sun, B., Huang, R.: Beyond talking–generating holistic 3d human dyadic motion for communication. International Journal of Computer Vision 133(5), 2910–2926 (2025)
  58. [58]Tevet, G., Raab, S., Cohan, S., Reda, D., Luo, Z., Peng, X.B., Bermano, A.H., van de Panne, M.: Closd: Closing the loop between simulation and diffusion for multi-task character control. In: International Conference on Learning Representations (ICLR) (2025)
  59. [59]Tevet, G., Raab, S., Gordon, B., Shafir, Y., Cohen-Or, D., Bermano, A.H.: Human motion diffusion model. arXiv preprint arXiv:2209.14916 (2022)arxiv.org/abs/2209.14916
  60. [60]Tseng, J., Castellon, R., Liu, K.: Editable dance generation from music. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 448–458 (2023)
  61. [61]Wang, H., Zhu, W., Miaoli, L., Xu, Y., Gao, F., Tian, Q., Wang, Y.: Aligning human motion generation with human perceptions. In: International Conference on Learning Representations (ICLR). vol. 2025, pp. 77626–77656 (2025)
  62. [62]Wang, Y., Wang, S., Zhang, J., Fan, K., Wu, J., Xue, Z., Liu, Y.: Timotion: Temporal and interactive framework for efficient human-human motion generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 7169–7178 (2025)
  63. [63]Wang, Y., Li, M., Liu, J., Leng, Z., Li, F.W., Zhang, Z., Liang, X.: Fg-t2m++: Llms-augmented fine-grained text driven human motion generation. International Journal of Computer Vision pp. 1–17 (2025)
  64. [64]Wang, Z., Chen, Y., Jia, B., Li, P., Zhang, J., Zhang, J., Liu, T., Zhu, Y., Liang, W., Huang, S.: Move as you say interact as you can: Language-guided human motion generation with scene affordance. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 433–444 (2024)
  65. [65]Wang, Z., Wang, J., Lin, D., Dai, B.: Intercontrol: Generate human motion interactions by controlling every joint. arXiv preprint arXiv:2311.15864 (2023)
  66. [66]Xie, Y., Jampani, V., Zhong, L., Sun, D., Jiang, H.: Omnicontrol: Control any joint at any time for human motion generation. arXiv preprint arXiv:2310.08580 (2023)arxiv.org/abs/2310.08580
  67. [67]Xing, J., Xia, M., Zhang, Y., Cun, X., Wang, J., Wong, T.T.: Codetalker: Speech-driven 3d facial animation with discrete motion prior. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 12780–12790 (2023)
  68. [68]Xu, Z., Lin, Y., Han, H., Yang, S., Li, R., Zhang, Y., Li, X.: Mambatalk: Efficient holistic gesture synthesis with selective state space models. In: Advances in Neural Information Processing Systems (NeurIPS) (2024)
  69. [69]Yang, S., Wu, Z., Li, M., Zhang, Z., Hao, L., Bao, W., Cheng, M., Xiao, L.: Diffusestylgesture: Stylized audio-driven co-speech gesture generation with diffusion models. In: Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI-23). pp. 5860–5868 (2023)
  70. [70]Yang, S., Xu, Z., Xue, H., Cheng, Y., Huang, S., Gong, M., Wu, Z.: Freetalker: Controllable speech and text-driven gesture generation based on diffusion models for enhanced speaker naturalness. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 7945–7949. IEEE (2024)
  71. [71]Yi, H., Liang, H., Liu, Y., Cao, Q., Wen, Y., Bolkart, T., Tao, D., Black, M.J.: Generating holistic 3d human motion from speech. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 469–480 (2023)
  72. [72]Yi, H., Thies, J., Black, M.J., Peng, X.B., Rempe, D.: Generating human interaction motions in scenes with text control. arXiv preprint arXiv:2404.10685 (2024)arxiv.org/abs/2404.10685
  73. [73]Yoon, Y., Cha, B., Lee, J.H., Jang, M., Lee, J., Kim, J., Lee, G.: Speech gesture generation from the trimodal context of text, audio, and speaker identity. ACM Transactions on Graphics 39(6), 1–16 (2020)
  74. [74]Zhang, J., Zhang, Y., Cun, X., Zhang, Y., Zhao, H., Lu, H., Shen, X., Shan, Y.: Generating human motion from textual descriptions with discrete representations. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 14730–14740 (2023)
  75. [75]Zhang, M., Cai, Z., Pan, L., Hong, F., Guo, X., Yang, L., Liu, Z.: Motiondiffuse: Text-driven human motion generation with diffusion model. arXiv preprint arXiv:2208.15001 (2022)arxiv.org/abs/2208.15001
  76. [76]Zhang, M., Jin, D., Gu, C., Hong, F., Cai, Z., Huang, J., Zhang, C., Guo, X., Yang, L., He, Y., et al.: Large motion model for unified multi-modal motion generation. In: European Conference on Computer Vision (ECCV). pp. 397–421. Springer (2024)
  77. [77]Zhang, M., Li, H., Cai, Z., Ren, J., Yang, L., Liu, Z.: Finemogen: Fine-grained spatio-temporal motion generation and editing. Advances in Neural Information Processing Systems (NeurIPS) 36, 13981–13992 (2023)
  78. [78]Zhang, P., Liu, P., Kim, H., Garrido, P., Chaudhuri, B.: Kinmo: Kinematic-aware human motion understanding and generation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2025)
  79. [79]Zhang, Q., Song, J., Huang, X., Chen, Y., Liu, M.Y.: Diffucollage: Parallel generation of large content with diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 10188–10198. IEEE (2023)
  80. [80]Zhang, X., Li, J., Zhang, J., Dang, Z., Ren, J., Bo, L., Tu, Z.: Semtalk: Holistic co-speech motion generation with frame-level semantic emphasis. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2025)
  81. [81]Zhang, Z., Wang, Y., Mao, W., Li, D., Zhao, R., Wu, B., Song, Z., Zhuang, B., Reid, I., Hartley, R.: Motion anything: Any to motion generation. arXiv preprint arXiv:2503.06955 (2025)arxiv.org/abs/2503.06955
  82. [82]Zhou, Z., Wan, Y., Wang, B.: A unified framework for multimodal, multi-part human motion synthesis. ArXiv abs/2311.16471 (2023), https://api.semanticscholar.org/CorpusID:265466120arxiv.org/abs/2311.16471
  83. [83]Zhou, Z., Wang, B.: Ude: A unified driving engine for human motion generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 5632–5641 (2023)
  84. [84]Zhu, L., Liu, X., Liu, X., Qian, R., Liu, Z., Yu, L.: Taming diffusion models for audio-driven co-speech gesture generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 10544–10555 (2023)

Paper details

Contents