Senior Site Reliability Engineer at Nebius responsible for ensuring reliability, performance, and observability of the Token Factory inference platform serving massive GPU workloads. The ideal candidate designs telemetry systems, optimizes Kubernetes infrastructure, implements resilience patterns, and drives incident response and post-mortem culture.