Senior Inference SW Engineer

Hewlett Packard Enterprise
US - North Carolina - Durham
View Company Profile / << Go Back

  • Job Type: Full time
  • Yesterday

Job Description

Design, implement, and operate major components of the LLM serving runtime and its Kubernetes orchestration layer, improving latency, throughput, GPU utilization, and distributed execution. Evaluate emerging inference technologies, resolve customer issues, and contribute through code reviews, mentoring, and strong engineering practices.

Requirements: Requires at least 8 years of software engineering experience, including 1–2 or more years working directly on LLM inference runtimes or production model serving; a degree in Computer Science or a related field is also specified. Candidates should have strong knowledge of inference engines and internals, Kubernetes, Go and Python, and be able to debug and profile C++/CUDA workloads.

Key Skills: LLM Inference, Inference Runtime Engineering, vLLM, SGLang, TensorRT-LLM, Continuous Batching, KV Cache Management, Kubernetes, Go, Python, C++, CUDA, Tensor Parallelism, Pipeline Parallelism, NCCL, GPU Profiling

Benefits: Health And Wellbeing Benefits, Personal And Professional Development




Fast Track Upload