Cloudjobs
US - New York - New York
View Company Profile /
<< Go Back
Design, implement, and operate infrastructure systems for model deployment and training on a large-scale GPU fleet. This includes managing job scheduling, cluster provisioning, and high-performance snapshot delivery.
Requirements: Candidates should have strong programming skills and experience with hyperscale compute systems, specifically using Azure and Kubernetes. An understanding of AI/ML workloads is considered a bonus.
Key Skills: Hyperscale compute systems, Azure, Kubernetes, Job scheduling, Cluster management, CI/CD systems, Programming, AI/ML workloads
Benefits: Relocation assistance
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX