Alignment Red Team - Research Engineer/Research Scientist

London, United Kingdom
Today
Posted
15 Sep 2026 (Today)

About the AI Security Institute

The AI Security Institute is the world's largest and best-funded team dedicated to understanding advanced AI risks and translating that knowledge into action. We’re in the heart of the UK government with direct lines to No. 10 (the Prime Minister's office), and we work with frontier developers and governments globally.

We’re here because governments are critical for advanced AI going well, and UK AISI is uniquely positioned to mobilise them. With our resources, unique agility and international influence, this is the best place to shape both AI development and government action.

The deadline for applying to this role is 11th October 2026, end of day, anywhere on Earth.

Team Description

Risks from misaligned AI systems are growing increasingly important as AI systems become more capable, autonomous, and integrated into society. Understanding these risks and stress-testing mitigations is crucial to ensuring advanced AI systems are developed and deployed safely and beneficially in the future.

The Alignment Red Team is a specialised subteam within AISI's wider Red Team focused on detecting and evaluating misalignment in frontier AI systems. We perform novel research to develop techniques for finding misalignment, and pre- and post-deployment evaluations of frontier AI systems to understand loss-of-control risks associated with models, such as deceptive alignment, research sabotage, and reward-seeking. We share our findings with frontier AI companies and the UK and allied governments, to inform their respective deployments, research, and policy-making. We also work directly with safety teams at frontier labs, sharing our evaluation findings to help improve their model alignment training and monitoring methodology.

We have conducted pre-deployment testing with multiple frontier AI companies for propensities related to research sabotage, cheating and unsanctioned cyber-attacks. We previously found that Claude models would sometimes refuse to help with benign AI safety research, an issue which Anthropic then evaluated for and fixed in their next model release.

About the Role

We're seeking Research Engineers and Research Scientists to join our Alignment Red Team. We are open to hires at junior, senior, staff and principal research scientist/engineer levels.

What You'll Be Doing

  • Researching methods to automatically search for misalignment in frontier models, including misalignment related to loss-of-control risks such as research sabotage and reward-seeking.
  • Building and running alignment evaluations relevant for loss-of-control risks that current benchmarks don’t capture.
  • Running pre-deployment evaluations to test the alignment of AI systems, and analysing and reporting results to frontier AI companies and UK and allied governments.
  • Contributing to public-facing research publications (like our published alignment evaluation case study) and technical reports that advance the field's understanding of misalignment risks and alignment evaluation methodology.
  • Designing and building software and tooling, including open-source software, for better alignment evaluations, improving efficiency, realism, and usability.

The work could also involve

  • Conducting threat modelling, analysis, and conceptual thinking to understand crucial model behaviours that could lead to loss of control (e.g. AI research assistants at frontier labs), translating abstract risk concepts into concrete, testable hypotheses.
  • Performing alignment incident investigations to understand after the fact what drove certain kinds of misaligned behaviour in frontier models.
  • Mentoring and advising external collaborators and researchers to do work relevant to the team’s goals and alignment testing more broadly.

Who we're looking for

Essential requirements

  • Ability to work autonomously on complex research projects involving substantial engineering. Have completed at least one significant research project in AI safety, security or alignment involving engineering, experiment design and analysis on frontier LLMs.
  • Strong software engineering and ML experience writing complex projects involving language models and ML, beyond just research code. 1+ years professional experience programming in Python for ML or SWE work.
  • Experience writing clean, documented research code for machine learning experiments, including experience with ML frameworks like PyTorch or evaluation frameworks like Inspect.
  • Proven ability in a team environment – flexible, adaptive to needs, and willing to contribute wherever necessary.
  • Impact-driven mindset, motivated by doing the most important work rather than what's superficially impressive.
  • High velocity and high-quality bar for outputs.

Highly Desirable

We don't expect candidates to have all of these – they're additional signals that help us identify

Related Jobs

View all jobs
Spotlight

Programme Manager (Forward Deployed)

M-1 Intelligence London, United Kingdom
£60,000 – £75,000 pa Remote

Research Engineer, Machine Learning (Reinforcement Learning)

Anthropic London, United Kingdom
£260,000 – £630,000 pa

Expression of Interest- Red Team

AI Security Institute London, United Kingdom
£65,000 – £145,000 pa On-site Clearance Required

[Expression of Interest] Research Engineer / Scientist, Alignment - London

Anthropic London, United Kingdom
Hybrid

Technical Director of AI Safety

Faculty AI London, United Kingdom
Hybrid Clearance Required

Industry Insights

Discover insightful articles, industry insights, expert tips, and curated resources.

What Is an AI Forward Deployed Engineer? The Fastest-Growing Job in AI for 2026

If you have been watching AI job boards over the past year, one title keeps surfacing again and again: the forward deployed engineer, or FDE. It has gone from a niche term known mainly to Palantir alumni to arguably the hottest role in the entire AI hiring market. Job postings for forward deployed engineers have exploded, salaries have climbed past levels most software engineers will ever see, and the biggest names in AI — OpenAI, Anthropic, Google, Salesforce, Databricks and Palantir — are all competing for the same small pool of talent. So what exactly is an AI forward deployed engineer, why has demand surged so dramatically, and how do you position yourself to land one of these roles? This guide breaks it all down for AI engineers, software engineers and data scientists looking at their next move.