Nebius seeks a Senior Site Reliability Engineer to own the reliability, performance, and observability of their Token Factory inference platform serving tens of thousands of GPUs. The ideal candidate will design telemetry pipelines, optimize Kubernetes infrastructure, and lead incident response and post-mortem culture for a high-scale AI cloud platform.