Jack & Jill
US - California - San Francisco
View Company Profile /
<< Go Back
The role involves designing and managing multi-region GPU deployments and EKS clusters to power real-time AI conversations. You will optimize CUDA hot paths and architect custom routing and scheduling logic to ensure sub-second latency globally.
Requirements: Candidates must have hands-on experience optimizing large-scale GPU inference workloads on providers like AWS or CoreWeave. Deep expertise in Kubernetes/EKS and a proven track record of senior technical leadership in infrastructure are required.
Key Skills: GPU Inference, Kubernetes, Amazon EKS, CUDA, Infrastructure Design, Cloud Deployment, Service Routing, Cluster Management, Performance Optimization, Technical Leadership
Benefits: Equity
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX