Site Reliability Engineer

Thought Machine
London, United Kingdom
2 months ago
Job Type
Permanent
Work Pattern
Full-time
Work Location
On-site
Seniority
Mid
Education
Degree
Posted
19 Mar 2026 (2 months ago)

Benefits

Employee share package High Glassdoor rating Fantastic workplace culture

Thought Machine’s mission is bold – to properly and permanently rid the world’s banks of legacy technology. To achieve this, we have developed the foundations of modern banking through core and payments technology which run natively in the cloud. What we are attempting is hard and means we need great people working together to build great technology.

We have grown rapidly in the past few years – growing our team to more than 550 individuals across offices in London, New York, Singapore and Sydney. We have raised more than $500m in funding and are now valued at $2.7bn. Our investors include Molten Ventures, Eurazeo, Intesa Sanpaolo, Temasek, Nyca Partners, JPMorgan Chase Strategic Investments, Standard Chartered Ventures, and more.

We have created a culture that enables our team to produce the best work in the industry while ensuring we have fun along the way. We're regularly cited as having a fantastic workplace culture and have been recognised by Sifted magazine as having one of the highest Glassdoor ratings for a UK fintech company and the industry's most generous employee share package. Named one of the world’s most innovative fintechs byGlobal Finance Magazine, we were also recognised by theFinancial Times as one of Europe’s fastest-growing companies for two consecutive years—and a UK Best Employer for 2026.

Thought Machine’s Site Reliability Engineers are the guardians of mission-critical systems for the world's most influential financial institutions. As a member of our elite, globally distributed team, you'll be entrusted with running and maintaining the robust production infrastructure that powers our customers' cutting-edge Core Banking and Payments platforms. This is an opportunity to make a tangible impact on the global financial landscape while collaborating with brilliant minds to solve complex engineering challenges.

This role will be part of the Site Reliability Engineering team at Thought Machine HQ in London. The team is deeply involved in tackling the technical challenges of executing Thought Machine’s growth ambitions - expect to be working with senior stakeholders in the organisation, our customers, and working on programmes and initiatives that are critical to the success of the company.

As an SRE at Thought Machine, you will be responsible for:

  • Supporting the product engineering teams in building highly fault-tolerant, scalable applications by participating in design discussions, engaging in RFCs and code reviews.

  • Contributing to the execution of department strategies such as implementing disaster recovery, backup, redundancy, and capacity planning activities.

  • Participating in a global on-call rotation responsible for identifying and fixing bottlenecks in SaaS customer environments.

  • Regular maintenance of production systems that host Vault products.

  • Contributing to the evolution of our SaaS products by building features that foster exceptional reliability and an unparalleled user experience.

  • Implementing and testing DR strategies to ensure the highest level of resilience and fault tolerance of the platform.

  • Maintaining high-quality written documentation of assets, processes and runbooks that are used by the team in their day-to-day operations.

  • Collaborating effectively with team members, actively participating in knowledge sharing, and continuously growing your own technical understanding of Vault Products.

What we’re looking for:

  • You have experience successfully delivering engineering tasks and projects with a focus on reliability and scalability.

  • You possess a good understanding of design patterns relevant to hosting and networking architectures.

  • You proactively champion product development, driven by a desire to build truly exceptional products, not just solve immediate challenges.

  • You have a strong background working in either Python, Golang or Java, having used one of these programming languages to build production level software.

  • You have experience working with Kubernetes or other container orchestration systems.

  • You have experience with automation/configuration management, e.g. Terraform, Puppet, Chef, Ansible.

  • You have a good understanding of one or more of the following areas: Database Administration, Networking, Observability Tools (such as Prometheus, Jaeger) or automation infrastructure.

  • You have solid experience working with either GCP or AWS.

Benefits:

  • Highly competitive salary

  • Pension plan (match up to 5%)

  • Life insurance - three times annual salary

  • Competitive maternity (six months fully paid) and paternity leave (four weeks fully paid)

  • Shared parental leave (matched to our maternity leave for the same point in time)

  • 25 days holiday and bank holidays

  • Flexible working hours

  • Cycle-to-work scheme

  • Electric car scheme

  • Season ticket loan

  • Access to outstanding learning materials and courses

  • Sports and hobby clubs, subsidised by Thought Machine

  • All the latest tech you need

  • Start the day properly with fresh fruit and cereals

  • Huge range of healthy (and not-so-healthy) snacks, smoothies and drinks

  • A talented and experienced team as your colleagues

  • An environment where we encourage learning and progress

  • Two charity days a year

  • Weekly food pop-up

We actively hire candidates who demonstrate technical excellence in their field and welcome people of all ages and backgrounds, providing everyone with equal access to professional development. You are encouraged to apply even if your experience doesn't accurately match the job description. We also encourage applications from those with different abilities, including candidates with ADHD, autism, dyslexia or dyspraxia.

Related Jobs

View all jobs
Spotlight

Forward Deployed Engineer

SolveAI London, United Kingdom
Hybrid
Spotlight

Machine Learning Engineer (Forward Deployed)

Mind Foundry Oxford/ Hybrid, Oxfordshire, United Kingdom

Senior Site Reliability Engineer

Thought Machine London, United Kingdom
Hybrid

Senior Cloud Site Reliability Engineer

Wayve London, United Kingdom
On-site

Director of Engineering (ML Platform), London

Isomorphic Labs London, United Kingdom
On-site

Global Head of Production Support

Thought Machine London, United Kingdom
Hybrid

Staff Cloud SRE – AI/ML Platform & GPU Compute

Wayve London, United Kingdom
On-site

Instructor in DevOps

Multiverse London, United Kingdom
Remote

Industry Insights

Discover insightful articles, industry insights, expert tips, and curated resources.

Where to Advertise AI Jobs in the UK (2026 Guide)

Where to advertise AI jobs UK in 2026: the specialist boards and communities that reach AI engineers, ML scientists and applied research talent in the UK. The candidate pool is small, highly informed and in demand across multiple sectors simultaneously. General job boards reach a broad audience but lack the specificity that AI professionals expect — and the filtering mechanisms they rely on. Specialist platforms, direct outreach and academic channels each serve a different part of the market. This guide, published by ArtificialIntelligenceJobs.co.uk, covers where to advertise AI roles in the UK in 2026, how the main platforms compare, what employers should expect to pay, and what the data says about time-to-hire across different role types.

AI Jobs UK 2026: What to Expect Over the Next 3 Years

AI Jobs UK 2026: roles, salaries and the generative AI, machine learning and applied AI hiring trends shaping UK artificial intelligence careers. Artificial intelligence is creating jobs faster than the market can name them. New roles are appearing every quarter, existing titles are splitting into specialisms, and the technologies underpinning it all are evolving at a pace that makes even last year's job descriptions feel dated. For job seekers, this presents a genuinely unusual challenge. In most industries, career planning means understanding a relatively stable landscape and working out where you fit within it. In AI, the landscape itself is being redrawn in real time. The roles with the most hiring activity in 2028 may not yet have a widely agreed job title in 2026. That's not a reason to feel overwhelmed — it's a reason to get informed. The candidates who thrive in this market aren't necessarily those with the longest CVs or the most credentials. They're the ones who understand the direction of travel: which skills are gaining value, which technologies are driving employer decisions, and how the definition of an "AI job" is expanding well beyond the tech sector. This article breaks down what the UK AI jobs market is likely to look like over the next three years — covering emerging job titles, the technologies reshaping hiring, the skills employers are prioritising, and how to position yourself ahead of the curve rather than behind it.