The Senior Site Reliability Engineer at Nebius plays a crucial role in maintaining the reliability and performance of the inference platform. This position involves designing telemetry systems, optimizing Kubernetes for GPU efficiency, and developing resilient infrastructure using Terraform. The ideal candidate should have extensive experience in site reliability engineering, particularly with distributed systems and GPU workloads.