OpenAI
US - California - San Francisco
View Company Profile /
<< Go Back
You will design, build, and operate Kubernetes-based controllers and distributed services to manage GPU compute infrastructure across sites. This includes developing lifecycle management for provisioning, configuration, and recovery of bare-metal systems to ensure high reliability and scalability.
Requirements: Candidates must have strong software engineering fundamentals and experience owning production distributed systems or infrastructure services. Proficiency with Kubernetes APIs, bare-metal node management, and designing reliable asynchronous workflows is essential.
Key Skills: Kubernetes, Distributed systems, Infrastructure as code, GPU compute, Bare-metal provisioning, API design, Linux, PXE, DHCP, DNS, BMCs, Firmware, Configuration management, System reliability, Concurrency, Reconciliation
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX