Crusoe
US - California - San Francisco
View Company Profile /
<< Go Back
Develop software for managing GPU server fleets and data center infrastructure, focusing on advanced diagnostics and observability. Create automation and AI agents for hardware remediation and critical environment management to maximize fleet availability.
Requirements: Requires software engineering experience with expertise in distributed systems, cloud platforms, and proficiency in languages like Go, Python, Java, or Rust. Candidates should be able to independently develop scalable solutions and set technical direction for complex projects.
Key Skills: Distributed Systems, Kubernetes, Infrastructure as Code, Google Cloud Platform, Go, Python, Java, Rust, GPU Fleet Management, Hardware Diagnostics, Observability Tooling, Automation, Temporal, NVIDIA NCCL, Pytorch, Reliability Engineering
Benefits: Restricted Stock Units, Health insurance (HDHP and PPO), Vision insurance, Dental insurance, HSA account contributions, Paid Parental Leave, Paid life insurance, Short-term disability, Long-term disability, Teladoc, 401(k) with 100% match up to 4%, Paid time off, Holiday schedule, Cell phone reimbursement, Tuition reimbursement, Calm app subscription, MetLife Legal, Commuter benefit
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX