Nebius seeks a Senior Site Reliability Engineer to own the reliability, performance, and observability of their Token Factory inference platform serving tens of thousands of GPUs. The ideal candidate will design telemetry systems, optimize Kubernetes infrastructure, and build automation to detect and remediate incidents across a massive-scale AI cloud platform.