Senior Researcher in Interpretability and AI Safety
This role involves researching interpretability, evaluations, and AI safety in continuously learning systems, with a focus on monitoring how capabilities and safety properties evolve over time. The researcher will conduct large-scale experiments on foundation models and collaborate within the Technical AI Governance programme. Work will include mechanistic analysis and developing benchmarks for dynamic AI systems.