Back to jobs

Senior Systems Software Engineer – GPU Performance at Scale

Full-time US, CA, Santa Clara On-site Posted 14 horas ago

NVIDIA

Opportunity published from Workday · applications stay with the original hiring source.

Source verified 6 hours agoThe original job page was active during UpVagas' latest lifecycle check.

About this opportunity

About the opportunity: NVIDIA is seeking a Senior Systems Software Engineer specializing in GPU Performance at Scale.

In this role, you will drive innovation in AI and GPU computing by contributing to world-class computing hardware and software.

You will provide insights on large-scale system composition and tuning mechanisms for high-performance compute runs while collaborating with researchers, developers, and customers to craft improved workflows and leading solutions.

You will engage with HPC, OS, CPU, GPU compute, and systems specialists to architect, build, and optimize large-scale performance platforms.

Key responsibilities

  • Lead the implementation of performance practices in large-scale GPU infrastructure, delivering powerful tools, methodologies, and flows to validate and improve multiple datacenter products concurrently.
  • Align next-generation AI workloads with next-generation datacenter builds for NVIDIA GPUs, CPUs, and networking hardware, engaging early with internal and customer teams across HW, FW, SW, and platforms.
  • Develop engineering solutions that provide continuous insights into the performance of AI workloads in evolving environments, generating swift insights into improvements and regressions.
  • Decompose high-complexity performance or stability issues into minimal reproduction cases, working towards identifying the root cause.
  • Participate in collaborations with various SW and FW teams including BMC, SBIOS, OS, and drivers to develop outstanding methods and tools.
  • Analyze, debug, and resolve critical firmware and software issues to achieve the highest AI workload performance at scale.

Qualifications

  • Proven understanding of accelerated computing software stacks (CUDA).
  • Experience with modern cloud and container-based enterprise computing architectures, with Slurm preferred.
  • Strong programming and scripting experience in C, C++, Python, and Bash.
  • Deep expertise in systems architecture and the impact of various components on performance.
  • Experience with container technology and Linux-based operating systems, with Docker preferred.
  • Experience supporting high-performance computing or deep learning in engineering or academic research communities.
  • Strong teamwork and communication skills, coupled with results-focused analytical abilities.
  • BS in Engineering, Mathematics, Physics, or Computer Science (or equivalent experience); MS or PhD desirable with 8+ years of applicable experience.

Nice to have

  • End-to-end GPU performance engineering from the profiler to systems analysis.
  • Linux systems programming and optimization experience.
  • Exposure to virtualization techniques and cloud platform solutions.
  • Experience with scheduling and resource management systems.
  • Experience with large-scale HPC environments.

Benefits

  • Base salary determined by location, experience, and peer pay.
  • Equity eligibility.
  • Comprehensive benefits package.

Job details

  • Base salary range: 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.
  • Applications accepted at least until October 9, 2026.
  • Posting type: Existing vacancy.

About the company: NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years through great technology and amazing people.

Today, the company is tapping into the unlimited potential of AI to define the next era of computing, where GPUs act as the brains of computers, robots, and self-driving cars.

Additional information

  • NVIDIA uses AI tools in its recruiting processes.

Application: UpVagas lists this opportunity from the hiring company or its recruitment platform. The application is completed at the original source.

Full listing

UpVagas condensed repetitive or legal boilerplate to make this page easier to read. The original hiring source contains the complete job description.

Source transparency. UpVagas organizes information provided by the company or recruitment source. Always review the original listing before submitting personal information.
Get new jobs on Telegram Tap the mascot to set your job alerts.
UpVagas mascot with Telegram