Senior Software Engineer, Quantized Inference

NVIDIA AI
US - Washington - Redmond
View Company Profile / << Go Back

  • Job Type: Full time
  • 30+ days ago

Job Description

Implement quantized and sparse recipes within inference engines and manage model export pipelines to ensure correct serialization. Develop benchmarking harnesses, data analysis tools, and improve developer productivity through infrastructure and CI improvements.

Requirements: Requires proficiency in Python and familiarity with C++, along with strong software engineering fundamentals and experience with ML accelerators. Candidates should have experience with PyTorch internals and a minimum of 4 years in a relevant software engineering role, preferably with a MS/PhD in Computer Science.

Key Skills: Python, C++, Triton Kernels, PyTorch, Quantized Inference, Model Compression, Machine Learning Accelerators, vLLM, TRT-LLM, SGLang, Megatron-LM, ModelOpt, Software Engineering, Data Analysis, Numerical Debugging, Large Language Models

Benefits: Equity, Benefits




Fast Track Upload