Xu Gui, Hua Yu, Dongsheng Zhou, Yaqing Hou

2026IEEE TRANSACTIONS ON EVOLUTIONARY COMPUTATION

DOI: 10.1109/tevc.2026.3696390

Abstract

Text-driven 3D human motion generation aims to synthesize diverse motion sequences that are semantically aligned with textual descriptions. Most existing generative approaches encode the entire sentence into a single semantic embedding, which inevitably compromises the semantic consistency of the generated motions. Furthermore, in long sequences, the lack of explicit physical constraints often leads to physically implausible motions, particularly in the later segments. To address these limitations, we propose LMPGen, a novel framework that formulates text-driven 3D human motion generation as a multiobjective optimization problem. Without introducing additional parameters, LMPGen enhances both the diversity and quality of the generated motions. Leveraging the semantic understanding capabilities of a large language model (LLM), we slice the input text into fine-grained action segments, each guiding the construction of semantically grounded objective functions. To effectively balance conflicting objectives, we adopt an evolutionary algorithm (EA) augmented by a pretrained generative model to initialize and guide population evolution. Each action segment serves as a semantic anchor, steering the search process toward diverse yet semantically faithful motion sequences. Furthermore, we incorporate physical criterion into the optimization process to ensure that the generated motions are smooth, stable, and physically plausible. Experimental results on mainstream human motion datasets demonstrate that our method outperforms existing approaches in terms of both semantic consistency and diversity, providing a robust solution for generating high-quality, text-aligned human motion sequences.

Citation format

GUI, Xu, et al. LLM-Guided multiobjective optimization for physically-plausible and diverse 3d human motion generation. IEEE TRANSACTIONS ON EVOLUTIONARY COMPUTATION, 2026.