Staff Software Engineer- Foundation Model Inference

DevHub
US - California - San Francisco
View Company Profile / << Go Back

  • Job Type: Full time
  • 22 days ago

Job Description

Build and optimize LLM infrastructure to power large-scale inference workloads for both partner and self-hosted models. Collaborate with cross-functional teams to improve the reliability, latency, and efficiency of distributed AI workloads.

Requirements: Requires over 8 years of experience in backend or infrastructure engineering with expertise in distributed systems and scalable APIs. Experience with ML infrastructure, GPU orchestration, and service-oriented architecture is essential.

Key Skills: Backend Engineering, Infrastructure Engineering, Distributed Systems, Scalable APIs, Cloud-native Infrastructure, Real-time Serving, ML Infrastructure, GPU Orchestration, Service-oriented Architecture, Deployment Pipelines, System Observability, LLM Infrastructure, Model Inference, PyTorch, vLLM

Benefits: Annual performance bonus, Equity




Fast Track Upload