Senior Site Reliability Engineer role at Nebius focused on ensuring reliability, performance, and observability of the Token Factory inference platform serving tens of thousands of GPUs. The ideal candidate will design telemetry pipelines, optimize Kubernetes infrastructure, and own incident response for a large-scale AI cloud platform.