Hewlett Packard Enterprise
US - Colorado - Fort Collins
View Company Profile /
<< Go Back
Design, implement, and operate major components of an enterprise LLM serving runtime, including engine integration, batching, KV cache management, quantized execution, and distributed GPU execution. Improve inference latency and throughput, contribute to Kubernetes orchestration, resolve customer issues, and provide technical leadership through reviews and mentoring.
Requirements: Requires at least eight years of software engineering experience, including one to two or more years working directly on LLM inference runtimes or production model serving. Candidates should have strong knowledge of inference engines and internals, Kubernetes architecture, Go and Python, and the ability to debug and profile C++/CUDA workloads; a computer science or related degree is specified.
Key Skills: LLM Inference, Inference Runtime Engineering, vLLM, SGLang, TensorRT-LLM, Continuous Batching, KV Cache Management, Quantization, Speculative Decoding, Tensor And Pipeline Parallelism, NCCL, Kubernetes, Go, Python, C++, CUDA
Benefits: Health And Wellbeing Benefits, Personal And Professional Development
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX