Nebius seeks a Senior Site Reliability Engineer to own the reliability, performance, and observability of their Token Factory inference platform serving tens of thousands of GPUs. The ideal candidate will design telemetry pipelines, optimize Kubernetes autoscaling, manage infrastructure-as-code, and lead incident response and post-mortem processes for a massive-scale AI inference system.