Hewlett Packard Enterprise
US - North Carolina - Durham
View Company Profile /
<< Go Back
Design, implement, and operate major components of the LLM serving runtime and its Kubernetes orchestration layer, improving latency, throughput, GPU utilization, and distributed execution. Evaluate emerging inference technologies, resolve customer issues, and contribute through code reviews, mentoring, and strong engineering practices.
Requirements: Requires at least 8 years of software engineering experience, including 1–2 or more years working directly on LLM inference runtimes or production model serving; a degree in Computer Science or a related field is also specified. Candidates should have strong knowledge of inference engines and internals, Kubernetes, Go and Python, and be able to debug and profile C++/CUDA workloads.
Key Skills: LLM Inference, Inference Runtime Engineering, vLLM, SGLang, TensorRT-LLM, Continuous Batching, KV Cache Management, Kubernetes, Go, Python, C++, CUDA, Tensor Parallelism, Pipeline Parallelism, NCCL, GPU Profiling
Benefits: Health And Wellbeing Benefits, Personal And Professional Development
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX