NVIDIA AI
US - Washington - Seattle
View Company Profile /
<< Go Back
The role involves optimizing LLM inference frameworks and model architectures for NVIDIA edge AI hardware. Responsibilities include developing inference recipes, characterizing multi-node behavior, and managing model validation workflows.
Requirements: Requires a degree in Computer Science or Engineering with over 12 years of experience in GPU computing and ML systems. Candidates must have strong C++/Python skills and hands-on experience with CUDA kernel development and container engineering.
Key Skills: GPU Computing, ML Systems, High-performance Inference, Python, C++, CUDA, Triton, LLM Inference, KV-cache Management, Continuous Batching, Quantization, Tensor Parallelism, Docker, OCI, NVIDIA Container Toolkit, Performance Analysis
Benefits: Equity, Comprehensive Benefits Package
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX