Music Technology and Sound StudiesSpeech and Audio ProcessingMusic and Audio Processing

Ge Zhu, Yutong Wen, Zhiyao Duan

2026.4.1Foundations and Trends in Signal Processing

DOI: 10.1108/ftsig-03-2026-140

Abstract

Diffusion models have emerged as powerful deep generative techniques, producing high-quality and diverse samples in applications in various domains, including audio. While existing reviews provide overviews, there remains limited in-depth discussion of these specific design choices. The audio diffusion model literature also lacks principled guidance for the implementation of these design choices and their comparisons for different applications. This survey provides a comprehensive review of diffusion model design with an emphasis on design principles for quality improvement and conditioning for audio applications. The authors adopt the score modeling perspective as a unifying framework that accommodates various interpretations, including recent approaches like flow matching. They systematically examine the training and sampling procedures of diffusion models and audio applications through different conditioning mechanisms. To provide an integrated, unified codebase and to promote reproducible research and rapid prototyping, they introduce an open-source codebase (Link to the website of github) that implements their reviewed framework for various audio applications. They demonstrate its capabilities through three case studies: audio generation, speech enhancement and text-to-speech synthesis, with benchmark evaluations on standard data sets.

Citation format

ZHU, Ge; WEN, Yutong; DUAN, Zhiyao. Audio generation through score-based generative modeling: Design principles and implementation. Foundations and Trends in Signal Processing, 2026, 20(1): 1–82.