Site Reliability Engineer (SRE) / Platform Engineer

London, City And County Of the City Of London, United Kingdom
Today
Job Type
Contract
Work Pattern
Full-time
Work Location
Hybrid
Seniority
Mid
Posted
31 Jul 2026 (Today)

Benefits

2 days onsite per week Potential extension opportunities

Site Reliability Engineer (GCP & Kubernetes)

London (Hybrid, 2 days onsite per week)

6-Month Contract | Inside IR35 | Competitive Day Rate

We are looking for an experienced Site Reliability Engineer (SRE) to join a high-performing technology team responsible for building, securing, and operating cloud-native platforms that support innovative AI-driven applications.

This is a key role focused on cloud infrastructure development, platform reliability, security hardening, and operational excellence. You will work closely with software engineers, platform teams, and technical stakeholders to design, deploy, and maintain resilient production systems capable of supporting complex and scalable workloads.

You will be responsible for the reliability, performance, scalability, and security of cloud-based applications and infrastructure, ensuring services are delivered efficiently and operate effectively in production environments.

Key Responsibilities:

Design, build, and maintain cloud infrastructure within Google Cloud Platform (GCP).

Deploy, manage, and optimise Kubernetes-based environments and containerised applications.

Develop and support backend services and operational tooling using Python.

Create and maintain Infrastructure as Code using Terraform.

Implement security hardening, monitoring, observability, and reliability best practices across cloud platforms.

Troubleshoot and resolve complex production issues, ensuring minimal service disruption.

Develop and improve CI/CD pipelines to enable efficient and reliable software delivery.

Work closely with development teams to improve platform resilience, scalability, and operational performance.

Implement monitoring, alerting, and incident response processes to support production systems.

Drive automation initiatives to reduce operational overhead and improve system reliability.We're Looking For:

Proven experience as a Site Reliability Engineer, Platform Engineer, Cloud Engineer, or DevOps Engineer.

Strong hands-on experience with Google Cloud Platform (GCP).

Experience deploying, operating, and supporting production cloud applications at scale.

Strong knowledge of Kubernetes and container orchestration technologies.

Commercial experience developing with Python.

Experience building and supporting backend services and APIs, ideally using FastAPI.

Strong knowledge of Terraform and Infrastructure as Code principles.

Experience designing highly available, resilient, and secure cloud environments.

Strong troubleshooting and incident management skills within production environments.

Experience with CI/CD pipelines, Git, GitHub, and DevOps best practices.

Strong understanding of cloud networking concepts, with particular emphasis on GCP networking.

Ability to take ownership of production systems and drive continuous improvement initiatives.Desirable:

Microsoft Azure experience.

Experience deploying and supporting AI/ML applications in production environments.

Exposure to Large Language Models (LLMs), NLP, or agent-based systems.

Experience with observability and monitoring tools such as Sentry, Prometheus, Grafana, or similar platforms.

Strong Docker and multi-container application architecture experience.

Experience working within scientific, pharmaceutical, genomics, or bioinformatics environments.

Start-up or high-growth technology company experience.

Open-source software contributions.

Technical leadership, mentoring, or coaching experience.Contract Details:

6-Month Initial Contract

Inside IR35

London-based

Hybrid Working (2 days onsite per week)

Competitive Day Rate

Potential extension opportunitiesThe Opportunity

This role offers the chance to work on modern cloud-native systems supporting innovative AI-enabled technologies. You'll be part of a collaborative engineering environment where reliability, automation, security, and scalability are key priorities, with the opportunity to make a significant impact on critical production platforms.

If you're a Site Reliability Engineer or Platform Engineer with strong GCP, Kubernetes, Terraform, and Python experience, we'd love to hear from you

Related Jobs

View all jobs
Spotlight

Solution Architect (Power Platform)

Loughborough University Loughborough, Leicestershire, United Kingdom
Hybrid
Spotlight

Senior AI Engineer

Bodyswaps London, United Kingdom
Hybrid

Site Reliability Engineer

Darktrace Cambridge, CB2 3BJ, United Kingdom
Hybrid

Senior Cloud Site Reliability Engineer

Wayve London, United Kingdom
On-site

Staff Software Engineer, AI Reliability Engineering

Anthropic London, United Kingdom

Software Engineer, GPU Infrastructure- ChatGPT Engineering

OpenAI London, United Kingdom
Hybrid

Platform Engineer (GCP)

HAYS Specialist Recruitment Manchester, United Kingdom
£50,000 – £65,000 pa Hybrid

Staff Cloud SRE – AI/ML Platform & GPU Compute

Wayve London, United Kingdom
On-site

Industry Insights

Discover insightful articles, industry insights, expert tips, and curated resources.

What Is an AI Forward Deployed Engineer? The Fastest-Growing Job in AI for 2026

If you have been watching AI job boards over the past year, one title keeps surfacing again and again: the forward deployed engineer, or FDE. It has gone from a niche term known mainly to Palantir alumni to arguably the hottest role in the entire AI hiring market. Job postings for forward deployed engineers have exploded, salaries have climbed past levels most software engineers will ever see, and the biggest names in AI — OpenAI, Anthropic, Google, Salesforce, Databricks and Palantir — are all competing for the same small pool of talent. So what exactly is an AI forward deployed engineer, why has demand surged so dramatically, and how do you position yourself to land one of these roles? This guide breaks it all down for AI engineers, software engineers and data scientists looking at their next move.