DevHub
US - California - San Francisco
View Company Profile /
<< Go Back
Build and optimize LLM infrastructure to power large-scale inference workloads for both partner and self-hosted models. Collaborate with cross-functional teams to improve the reliability, latency, and efficiency of distributed AI workloads.
Requirements: Requires over 8 years of experience in backend or infrastructure engineering with expertise in distributed systems and scalable APIs. Experience with ML infrastructure, GPU orchestration, and service-oriented architecture is essential.
Key Skills: Backend Engineering, Infrastructure Engineering, Distributed Systems, Scalable APIs, Cloud-native Infrastructure, Real-time Serving, ML Infrastructure, GPU Orchestration, Service-oriented Architecture, Deployment Pipelines, System Observability, LLM Infrastructure, Model Inference, PyTorch, vLLM
Benefits: Annual performance bonus, Equity
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX