The Senior Site Reliability Engineer role at Nebius is essential for ensuring the reliability and performance of the inference platform. This position involves designing telemetry pipelines, tuning Kubernetes for efficiency, and creating resilient infrastructure using Terraform. The ideal candidate should have extensive experience with production systems, particularly with GPU workloads, and possess a strong command of tools like Prometheus and Grafana.