Senior Cloud Site Reliability Engineer
This role involves building and scaling the reliability foundations of an AI cloud platform, including model development and GPU compute systems. The engineer will define SRE frameworks, ensure platform resilience, and enable scalable AI deployment through automation and observability. It's a founding role requiring deep expertise in distributed systems, Kubernetes, and cloud infrastructure.