Home
Experience
Projects
Publications
Teaching
Talks
AI safety
A million authors of alignment
Community-written narratives, folded in during mid-training, as a route to alignment - capturing human values through storytelling and giving communities a real say over open-weight model behaviour.
Diksha Gupta
LessWrong, 2026
Red-teaming LLM unlearning: LUNAR's "forgotten" knowledge is still recoverable
Two white-box attacks recovered knowledge that LUNAR, a state-of-the-art unlearning method, was supposed to have removed - suggesting localised edits may not be enough for robust unlearning.
Diksha Gupta
LessWrong, 2026
knowledge removal
Cite
×