Google DeepMind
I am a Research Scientist at Google DeepMind working on AGI readiness. My interests lie at the intersection of interpretability, safety, and security. I am exploring both black-box and mechanistic interpretability concurrently, to bridge the gap between observable behaviors and their underlying neural mechanisms. My current focus is on studying continual learning systems, particularly—but not limited to—emerging misbehaviors.
Previously, I worked at Meta on privacy mechanisms, federated learning, multimedia retrieval, and copyright protection.