Anthropic Fellows Program, AI Safety & Security
This role involves conducting full-time empirical AI safety research over a 4-month period, aligned with Anthropic’s priorities such as scalable oversight, adversarial robustness, and mechanistic interpretability. Fellows will work on a self-directed project with mentorship from senior researchers, aiming to produce a public output like a research paper. The program offers a structured environment with access to compute funding, a shared workspace, and integration into the broader AI safety research community.