NVIDIA AI
US - California - Santa Clara
View Company Profile /
<< Go Back
You will design and develop a massively distributed scalable platform to identify, diagnose, and remediate non-performant GPU assets within DGX Cloud. Additionally, you will collaborate across teams to ensure production AI clusters run reliably and consistently with maximum performance.
Requirements: Candidates must have 5+ years of experience in a similar software engineering role, specifically with large-scale production systems. Proficiency in React, TypeScript, Golang, and SQL databases is required, along with a BS in Computer Science or equivalent experience.
Key Skills: React, TypeScript, Golang, Kubernetes, PostgreSQL, Temporal, Bazel, GPU resource scheduling, Cluster operations, Node health monitoring, Distributed systems, Software engineering, Incident management, Scalable infrastructure, AI infrastructure
Benefits: Equity, Benefits
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX