Jack & Jill
US - California - San Francisco
View Company Profile /
<< Go Back
The role involves owning and managing multi-region GPU inference infrastructure to power real-time AI conversations with sub-second latency. Key tasks include architecting EKS clusters and collaborating with researchers to optimize CUDA hot paths.
Requirements: Candidates must have hands-on experience deploying large-scale GPU workloads on providers like AWS or CoreWeave. Deep expertise in Kubernetes/EKS and a proven track record of senior technical leadership are required.
Key Skills: GPU Inference, Kubernetes, Amazon EKS, CUDA, Infrastructure Design, Cloud Computing, Service Routing, Cluster Management
Benefits: Equity
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX