Towards General Audio Intelligence
COM3 Level 2
SR21, COM3 02-60

Abstract:
Audio-Visual General Intelligence—the capacity of AI agents to deeply understand and reason about all types of auditory and visual inputs, including speech, environmental sounds, and music, and to combine them with images and videos—is crucial for enabling AI to interact seamlessly and naturally with our world. Despite this importance, audio understanding has traditionally lagged advancements in vision and language processing. This gap stems from significant challenges, including limited datasets, the complexity of audio signals, and a shortage of advanced neural architectures and effective training methods explicitly tailored for audio.
In this talk, we overview our work on audio and audio-visual understanding, and the development of audio and audio-visual LLMs. We provide an overview of audio large language models (ALLMs) used for advanced audio perception and complex reasoning. This includes GAMA, which is built with a specialized architecture, optimized audio encoding, and a novel alignment dataset. We also introduce ReCLAP, a state-of-the-art audio-language encoder, and CompA, one of the first projects to tackle compositional reasoning in audio-language models—a critical challenge given the inherently compositional nature of audio. We also discuss the Audio Flamingo series and Music Flamingo, which are our open ALLM models that provide advanced long-audio understanding and reasoning capabilities across speech, sound, and music. Finally, we present audio-visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed to understand and reason over long, complex real-world (audio-visual) videos, and we highlight its performance across different scenarios.
Bio:
Dinesh Manocha is the Paul Chrisman-Iribe Chair in Computer Science & ECE and a Distinguished University Professor at the University of Maryland, College Park. His research interests include virtual environments, physically-based modeling, and robotics. His group has developed numerous software packages that are standard and licensed to over 60 commercial vendors. He has published more than 900 papers & supervised 69 PhD dissertations. His group has received more than 22 best paper and test-of-time awards at leading conferences in computer graphics, solid modeling, multimedia, VR, and robotics. Manocha is a Fellow of AAAI, AAAS, ACM, IEEE, and NAI, a member of ACM SIGGRAPH and IEEE VR Academies, and a recipient of the Bézier Award from the Solid Modeling Association. He received the Distinguished Alumni Award from IIT Delhi and the Distinguished Career in Computer Science Award from the Washington Academy of Sciences. He co-founded Impulsonic, a physics-based audio simulation technology developer, which Valve Inc. acquired in November 2016.

