Towards Robust and Scalable Post Hoc Model Adjustment Techniques in Deep Learning Models
Amula Venkat Adithya
Abstract
Despite their high performance, deep learning models often require robust post-hoc adjustments; these interventions serve not just to correct inherent vulnerabilities or biases, but also to adapt the model to evolving requirements that emerge after training. Three main lines of work satisfy these criteria, namely: Machine Unlearning, Concept Based Model Improvement, and Backdoor Defense techniques. These techniques operate with different goals and at different levels of a data hierarchy: sample-level (individual data points), concept-level (abstract features or subgroups within classes), and class-level (entire categories)—each addressing undesirable behaviors at varying degrees of granularity. We aim to answer two main questions in this thesis. To what extent can a unified framework address the diverse requirements of post-hoc model correction? And does the level of granularity in the data hierarchy imply easier scalability to all corrective requirements? Our primary contribution is Prototype Guided Backdoor Defense (PGBD), a novel backdoor defense method. PGBD exploits displacements in the geometric spaces of activations to penalize movements towards the trigger. This is done using a novel sanitization loss of a post-hoc fine-tuning step. This approach scales to all types of attacks and trig- gers, and achieves better performance across settings. We also present the first defense against semantic attacks on a new celebrity face images dataset. As a secondary contribution, we conduct a systematic review of Machine Unlearning, establishing a novel taxonomy of its core ideas, diverse settings, and varied goals. Consequently, we consider the setting of Corrective Unlearning and explore the possibility of scaling Unlearning methods to the problem setting of Backdoor Defense. Results show that while Un- learning methods perform a good job of removing memorization of individual data points, they struggle to mitigate shortcuts learned through trigger features. Alternatively, leveraging PGBD’s configurability, we extend PGBD to the Corrective Unlearning setting and show that it outperforms all existing un- learning methods. Encouraged by these results, we argue that activation space manipulation provides a powerful and flexible blueprint for many post-hoc model adjustment tasks, particularly those involving concept-level or trigger-based vulnerabilities. While gradient-based unlearning remains the superior ap- proach for precise, sample-level forgetting, its application is often specialized. Our findings suggest that configurable, representation-agnostic interventions in activation space offer greater versatility across a broader spectrum of post-hoc corrections.
| Year of completion: | May 2026 |
| Advisor : |
Narayanan P J |
Related Publications
Downloads