Senior Site Reliability Engineer role at Nebius focused on ensuring reliability, performance, and observability of the Token Factory inference platform serving large-scale GPU deployments. The ideal candidate will design telemetry pipelines, optimize Kubernetes infrastructure, and lead incident response and post-mortem processes for a platform supporting foundation model inference at massive scale.