Senior Software Engineer - AI Inference Performance

NVIDIA AI
US - California - Santa Clara
View Company Profile / << Go Back

  • Job Type: Full time
  • 5 days ago

Job Description

Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput.

Requirements: Requires over 6 years of experience in full-stack AI inference performance with strong programming skills in Python, C++, or Rust and expertise in CUDA. A degree in Computer Science or Computer Engineering is required along with deep knowledge of GPU architecture and profiling tools.

Key Skills: CUDA, Python, C++, Rust, LLM Inference, VLM Inference, GPU Architecture, Nsight Systems, Nsight Compute, PyTorch Profiler, TensorRT-LLM, vLLM, Triton, CUTLASS, Distributed Systems, Quantization

Benefits: Equity, Generous Benefits Package




Fast Track Upload