Member of Technical Staff (Software Engineer, Inference & Training Platform)

Perplexity
US - California - San Francisco
View Company Profile / << Go Back

  • Job Type: Full time
  • 30+ days ago

Job Description

The role involves building a self-serve compute platform for running training and inference workloads while managing a GPU fleet. Responsibilities include operating the GPU fleet, solving for GPU scarcity, and ensuring the reliability of both training jobs and inference services.

Requirements: Candidates should have deep experience with Kubernetes, GPU clusters, and distributed systems fundamentals. Proficiency in programming languages such as Go, Rust, or C++ is also required, along with experience in managing workloads across multiple cloud providers.

Key Skills: Kubernetes, GPU Clusters, CUDA, Networking, Distributed Systems, Go, Rust, C++, Inference Services, Training Jobs, Fault Tolerance, Autoscaling, Observability, Scheduling, Resource Allocation, Multi-Cloud




Fast Track Upload