NVIDIA AI
US - California - Santa Clara
View Company Profile /
<< Go Back
Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput.
Requirements: Requires over 6 years of experience in full-stack AI inference performance with strong programming skills in Python, C++, or Rust and expertise in CUDA. A degree in Computer Science or Computer Engineering is required along with deep knowledge of GPU architecture and profiling tools.
Key Skills: CUDA, Python, C++, Rust, LLM Inference, VLM Inference, GPU Architecture, Nsight Systems, Nsight Compute, PyTorch Profiler, TensorRT-LLM, vLLM, Triton, CUTLASS, Distributed Systems, Quantization
Benefits: Equity, Generous Benefits Package
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX