Senior Site Reliability Engineer role at Nebius focused on ensuring reliability, performance, and observability of the Token Factory inference platform serving massive GPU workloads. The ideal candidate will design telemetry systems, optimize Kubernetes infrastructure, and establish incident response automation for an AI cloud platform handling tens of thousands of GPUs.