Member of Technical Staff (Software Engineer, Inference & Training Platform)

Perplexity
US - California - San Francisco
View Company Profile / << Go Back

  • Job Type: Full time
  • 30+ days ago

Job Description

The role involves building a self-serve compute platform for training and inference workloads while managing a GPU fleet across multiple cloud providers. The engineer will also ensure reliability, fault tolerance, and efficient resource allocation for both long-running training jobs and production inference services.

Requirements: Candidates should have deep experience with Kubernetes, GPU clusters, and distributed systems fundamentals. Proficiency in programming languages such as Go, Rust, or C++ is also required, along with experience in managing high-availability inference services.

Key Skills: Kubernetes, GPU Clusters, CUDA, Networking, Distributed Systems, Go, Rust, C++, Inference Services, Fault Tolerance, Autoscaling, Observability, Scheduling, Resource Allocation, High-Availability, Multi-Cloud




Fast Track Upload