Representation Structure in LLMs

Recent work has hypothesised that large language models may represent some concepts as multi-dimensional manifolds instead of single directions. This has encouraged new steering and probing methods that respect manifold geometry. These methods assume that a path along the manifold corresponds to continuous changes in the represented concept and resulting behaviour. If the geometry is instead discrete, such steering places activations off-distribution rather than capturing intermediate values. I am currently testing whether this assumption holds.
More details soon.
Talks
- Aug 2026 Pivotal Spotlight, London Initiative for Safe AI (LISA) - selected
- Aug 2026 Pivotal Lightning talk, UK AISI - selected