Senior Site Reliability Engineer responsible for ensuring reliability, performance, and observability of Nebius's Token Factory inference platform serving massive GPU cloud workloads. The ideal candidate designs telemetry pipelines, optimizes Kubernetes infrastructure, and builds automation systems to detect and prevent incidents across a high-scale AI deployment platform.