Hewlett Packard Enterprise
US - Texas - Spring
View Company Profile /
<< Go Back
Design and maintain major components of the LLM serving runtime, including engine integration, batching, KV cache management, quantized execution, and distributed execution. Improve inference performance and Kubernetes orchestration, evaluate emerging serving technologies, resolve customer issues, and mentor engineers.
Requirements: Requires at least eight years of software engineering experience, including one to two or more years working directly on LLM inference runtimes or production model serving, plus a degree in computer science or a related field. Candidates should have strong knowledge of inference engines and internals, Kubernetes architecture, Go and Python, and the ability to debug and profile C++/CUDA.
Key Skills: LLM Inference, Inference Runtime Engineering, vLLM, SGLang, TensorRT-LLM, Continuous Batching, KV Cache Management, Quantization, Speculative Decoding, Tensor and Pipeline Parallelism, NCCL, Kubernetes, Go, Python, C++, CUDA
Benefits: Health and Wellbeing Benefits, Professional Development
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX