Member of Technical Staff (Software Engineer, Inference & Training Platform)

Perplexity
US - California - San Francisco
View Company Profile / << Go Back

  • Job Type: Full time
  • 30+ days ago

Job Description

Build and operate a self-serve compute platform to manage GPU clusters across multiple cloud providers for training and inference workloads. Design scheduling and placement logic to optimize GPU utilization and ensure high availability and fault tolerance.

Requirements: Requires deep expertise in Kubernetes custom operators, GPU hardware (NVIDIA, CUDA), and distributed systems programming in Go, Rust, or C++. Candidates should have experience managing large-scale compute across multiple clouds and supporting both training and inference services.

Key Skills: Kubernetes, GPU Cluster Management, Distributed Systems, Go, Rust, C++, CUDA, Multi-cloud Orchestration, Resource Allocation, Fault Tolerance, Inference Serving, vLLM, SGLang, TensorRT-LLM, InfiniBand, RoCE




Fast Track Upload