Senior Site Reliability Engineer role at Nebius focused on ensuring reliability, performance, and observability of the Token Factory inference platform serving massive GPU cloud infrastructure. The ideal candidate will design telemetry pipelines, optimize Kubernetes systems, manage infrastructure-as-code, and lead incident response and post-mortem practices for a high-scale AI inference platform.