Cisco
US - California - San Jose
View Company Profile /
<< Go Back
Run and analyze AI training, inference, and HPC workloads on GPU-accelerated systems, using telemetry and profiling tools to identify compute, memory, interconnect, storage, and scaling bottlenecks. Build automation for performance testing, regression detection, and failure triage, collaborate with hardware and software teams to resolve issues, and document procedures and results.
Requirements: Requires a bachelor’s degree and at least five years of related experience, a master’s degree and at least three years, a PhD with relevant research experience, or equivalent work experience. Candidates need Linux systems debugging, testing, or tuning experience; experience evaluating AI, machine learning, or HPC workloads; and at least three years with a programming or scripting language, while GPU technologies, profiling tools, distributed workload environments, and GPU system architecture are preferred.
Key Skills: Linux Systems Debugging, AI Workload Performance Analysis, HPC Workloads, Python, C, C++, Go, Bash, GPU Computing, CUDA, ROCm, NCCL, RCCL, Performance Profiling, Kubernetes, Slurm
Benefits: Medical Insurance, Dental Insurance, Vision Insurance, 401(k) Matching, Paid Parental Leave, Short-Term Disability Coverage, Long-Term Disability Coverage, Life Insurance, Restricted Stock Units, Paid Holidays, Floating Holiday, Birthday Leave, Year-End Holiday Shutdown, Personal Wellness Days, Paid Vacation, Sick Leave, Family Emergency Leave, Paid Volunteer Days, Annual Bonus
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX