Senior Inference SW Engineer

Hewlett Packard Enterprise
US - Colorado - Fort Collins
View Company Profile / << Go Back

  • Job Type: Full time
  • Yesterday

Job Description

Design, implement, and operate major components of an enterprise LLM serving runtime, including engine integration, batching, KV cache management, quantized execution, and distributed GPU execution. Improve inference latency and throughput, contribute to Kubernetes orchestration, resolve customer issues, and provide technical leadership through reviews and mentoring.

Requirements: Requires at least eight years of software engineering experience, including one to two or more years working directly on LLM inference runtimes or production model serving. Candidates should have strong knowledge of inference engines and internals, Kubernetes architecture, Go and Python, and the ability to debug and profile C++/CUDA workloads; a computer science or related degree is specified.

Key Skills: LLM Inference, Inference Runtime Engineering, vLLM, SGLang, TensorRT-LLM, Continuous Batching, KV Cache Management, Quantization, Speculative Decoding, Tensor And Pipeline Parallelism, NCCL, Kubernetes, Go, Python, C++, CUDA

Benefits: Health And Wellbeing Benefits, Personal And Professional Development




Fast Track Upload