Skip to main navigation Skip to search Skip to main content

Interpretable timbre synthesis using variational autoencoders regularized on timbre descriptors

  • Technological University Dublin

Research output: Chapter in Book/Report/Conference proceedingsConference proceedingpeer-review

Abstract

Controllable timbre synthesis has been a subject of research for several decades, and deep neural networks have been the most successful in this area. Deep generative models such as Variational Autoencoders (VAEs) have the ability to generate a high-level representation of audio while providing a structured latent space. Despite their advantages, the interpretability of these latent spaces in terms of human perception is often limited. To address this limitation and enhance the control over timbre generation, we propose a regularized VAE-based latent space that incorporates timbre descriptors. Moreover, we suggest a more concise representation of sound by utilizing its harmonic content, in order to minimize the dimensionality of the latent space.
Original languageEnglish
Title of host publicationProceedings of the 26th International Conference on Digital Audio Effects (DAFx23), Copenhagen, Denmark, 4 - 7 September 2023
DOIs
Publication statusPublished - 7 Sept 2023
Externally publishedYes
Event26th International Conference on Digital Audio Effects (DAFx23) - Copenhagen, Denmark
Duration: 4 Sept 20237 Sept 2023
https://www.dafx.de/

Publication series

NameProceedings of the International Conference on Digital Audio Effects, DAFx
PublisherDAFx
ISSN (Print)2413-6700

Conference

Conference26th International Conference on Digital Audio Effects (DAFx23)
Country/TerritoryDenmark
CityCopenhagen
Period4/09/237/09/23
Internet address

Fingerprint

Dive into the research topics of 'Interpretable timbre synthesis using variational autoencoders regularized on timbre descriptors'. Together they form a unique fingerprint.

Cite this