Jack & Jill
US - California - San Francisco
View Company Profile /
<< Go Back
Design and manage multi-region GPU deployments and EKS clusters to power real-time AI conversations. Optimize CUDA hot paths and backend services to ensure sub-second latency globally.
Requirements: Requires hands-on experience optimizing large-scale GPU inference workloads on providers like AWS or CoreWeave. Deep expertise in Kubernetes and a proven track record of senior technical leadership are essential.
Key Skills: GPU Inference, Kubernetes, EKS, CUDA, Infrastructure Design, Cloud Deployment, Service Routing, Cluster Management
Benefits: Equity
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX