Senior Inference SW Engineer

Hewlett Packard Enterprise
US - Texas - Spring
View Company Profile / << Go Back

  • Job Type: Full time
  • Yesterday

Job Description

Design and maintain major components of the LLM serving runtime, including engine integration, batching, KV cache management, quantized execution, and distributed execution. Improve inference performance and Kubernetes orchestration, evaluate emerging serving technologies, resolve customer issues, and mentor engineers.

Requirements: Requires at least eight years of software engineering experience, including one to two or more years working directly on LLM inference runtimes or production model serving, plus a degree in computer science or a related field. Candidates should have strong knowledge of inference engines and internals, Kubernetes architecture, Go and Python, and the ability to debug and profile C++/CUDA.

Key Skills: LLM Inference, Inference Runtime Engineering, vLLM, SGLang, TensorRT-LLM, Continuous Batching, KV Cache Management, Quantization, Speculative Decoding, Tensor and Pipeline Parallelism, NCCL, Kubernetes, Go, Python, C++, CUDA

Benefits: Health and Wellbeing Benefits, Professional Development




Fast Track Upload