Music ControlNet: Multiple Time-Varying Controls for Music Generation
About this Topic:
By 2023, text-to-audio models were able to generate convincing music from a short text prompt. However, while a text prompt captures well the broad style and mood, it is a less suitable medium for the moment-to-moment choices a composer makes, such as where a beat falls, how a passage builds and releases, and which melodic line the music should follow. Music ControlNet was proposed to incorporate these fine-grained controls.
In this webinar, the presenters discuss how they revamped the pixel-wise control mechanism of ControlNet, developed for image generation, to enable time-varying control over melody, rhythm, and musical intensity. They model music as an image-like spectrogram generated by a diffusion process. Each control is aligned to the spectrogram in time but describes each frame only coarsely, such as by a single pitch class or intensity level, and the model thus learns how a control should map onto the frequencies and generates the rich spectral details that remain consistent with it. They further show that these controls are composable and flexible: a user can apply any subset of the controls the model was trained on, and can specify each control over only part of the duration, letting the model improvise the rest. They conclude the talk with recent efforts across the research community toward control, editing, and integration into existing music workflows, and offer commentary on future research directions.
About the Presenters:
Shih-Lun Wu received the B.Sc. degree computer science from National Taiwan University and the M.Sc. in computer science and language technologies from Carnegie Mellon University in 2021 & 2024 respectively. He is pursuing the Ph.D. in electrical engineering and computer science at Massachusetts Institute of Technology.
He is concurrently a Research Scientist on the Music AI team at Adobe Research, where he previously interned for two summers (2023 and 2025). From 2020 to 2022, he served as a Research Engineer at Taiwan AI Labs, working on music generation technologies. His long-term research interests span controllable music generation/editing, multimodal learning, and large model adaptation.
Mr. Wu was awarded the Siebel Scholarship (2024) for outstanding research and received the university-wide Best Bachelor's Thesis award from National Taiwan University (2021). He has published more than 15 peer-reviewed papers, many of them at premier IEEE SPS venues including T-ASLP, T-MM, and ICASSP, where he also served frequently as a reviewer.
Nicholas J. Bryan (SM’23) received the B.S. and B.A degrees in both music and electrical engineering from the University of Miami, graduating summa cum laude and the M.A., M.Sc. and Ph.D. Degrees in electrical engineering from Stanford University (CCRMA).
He is currently the Head of the Music AI team and Principal Scientist at Adobe Research, where he leads the development of music and generative AI, focusing on human-AI music co-creation. Most recently, he led the cross-team effort to ship Adobe’s Generate Soundtrack feature, a music generation tool designed to be commercially safe for any use.
Dr. Bryan is an Adobe Distinguished Inventor and a recipient of two Best Paper awards. To date, he has published 43 peer-reviewed papers and holds 14 patents, with over 20 more patents pending. He has also remains committed to the academic research community, having served as General Co-Chair of IEEE WASPAA 2023, two terms on the IEEE Audio and Acoustic Signal Processing Technical Committee, and as a member on five Ph.D. thesis committees. He also has been a musician since childhood, performed at Carnegie Hall, and has been on the front-page of the New York Times.
Want to learn more about upcoming events & webinars?
Visit the events section of the Signal Processing website to see all upcoming lectures, workshops, webinars, and more.