Hi, I am Diksha :)
I study how neural networks solve complex cognitive tasks, and diagnose when and how they fail. My interests lie at the intersection of LLM training dynamics, alignment, and interpretability.
I am currently a senior research fellow at Pivotal, working with Stefan Heimersheim on the continuity of LLM activation manifolds and what that implies for manifold-aware steering and probing. Previously, I worked on red-teaming unlearning methods for open-weight safety.
Before that, I was a senior research fellow at the Sainsbury Wellcome Centre at University College London, with the Behrens and Mrsic-Flogel labs, studying compositional representations and their learning algorithms in artificial and biological neural networks. During my PhD with Carlos Brody at Princeton University, I developed ablation-aware methods to train RNNs on cognitive tasks to improve capability attribution, and better models to track systematic biases in decision-making and predict previously unpredictable failures, working from neural recordings and causal experiments in rodents playing carefully crafted structured games.
LLMsStructure & attribution
Representation Structure in LLMs
Recent work has hypothesised that large language models may represent some concepts as multi-dimensional manifolds instead of single directions. This has encouraged new steering and probing methods that respect manifold geometry. …
LLMsFailure modes
Knowledge Removal in LLMs
A large portion of safety and alignment research has focused on closed frontier models, due to their advanced capabilities. However, the gap between open and closed models is narrowing fast, and open-weight models are now being …
BrainsStructure & attribution
Mechanistic Attribution in Decision Circuits
I am fascinated by how intelligent behavior emerges from the coordinated activity of networks of neurons. This page highlights my research aimed at uncovering the representations, architectures, and dynamics that are crucial for …
BrainsFailure modes
Failure Modes in Sequential Decisions
Neural activations are messy and complicated, so the behaviour of an agent often constrains a system’s computational model more tightly than measurements of its internals (Niv 2021). A large part of my research focuses on …