Investigating Auditory Concepts in Deep Neural Networks


Pratyaksh Gautam

Abstract

With the advent of deep learning [50], deep neural networks (DNNs) have shown a remarkable ability to perform complex tasks across a wide range of domains. These successes span language mod- eling [17, 29, 76], image classification [30, 48], and automatic speech transcription [9]. In the domain of non-speech audio specifically, DNNs have been deployed with great success to classify audio in- stances, identifying sounds, instruments, and musical genres [8, 31, 44]. Furthermore, multi-modal large language models (MLLMs) capable of processing both audio and text [24, 28, 37] have enabled capabilities such as captioning and question-answering for audio inputs. Beyond analysis, DNNs are now used to clone voices [5, 19] and generate music directly as raw waveforms [2, 25, 36, 53]. Collec- tively, these advances highlight the growing breadth and sophistication of deep learning in processing auditory information.

Despite their empirical success, the interpretability of DNNs remains a significant challenge; the internal mechanisms that give rise to model behavior are not yet fully understood. Because these models build understanding through many layers of non-linear transformations on high-dimensional data, it remains difficult to discern how their internal representations give rise to the behaviors we observe.

 

Year of completion:  June 2026
 Advisor :

Vinoo Alluri and Makarand Tapaswi


Related Publications


    Downloads

    thesis