Senior System Software Engineer – GPU Performance

NVIDIA
Opportunity published from Workday · applications stay with the original hiring source.
About the opportunity
NVIDIA is seeking a motivated Performance Engineer for the GPU Communications Libraries and Networking team.
You will influence the roadmap of communication libraries such as NCCL, NVSHMEM, and UCX, which power Deep Learning and HPC applications running at scales up to tens of thousands of GPUs.
This is an exceptional opportunity to advance the state of the art in high-performance computing and communication.
Key responsibilities
- Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters.
- Study the interaction of our libraries with all hardware (GPU, CPU, networking) and software components in the stack.
- Evaluate proof-of-concepts and conduct trade-off analysis when multiple solutions are available.
- Triage and root-cause performance issues reported by our customers.
- Collect performance data and build tools and infrastructure to visualize and analyze the information.
- Collaborate with a dynamic team across multiple time zones.
Qualifications
- M.S. (or equivalent experience) or PhD in Computer Science or a related field with relevant performance engineering and HPC experience.
- 3+ years of experience with parallel programming and at least one communication runtime (MPI, NCCL, UCX, NVSHMEM).
- Experience conducting performance benchmarking and triage on large-scale HPC clusters.
- Good understanding of computer system architecture, hardware-software interactions, and operating systems principles.
- Ability to implement micro-benchmarks in C/C++ and read or modify the codebase when required.
- Ability to debug performance issues across the entire hardware-software stack, with proficiency in a scripting language, preferably Python.
- Familiarity with containers, cloud provisioning, and scheduling tools (Kubernetes, SLURM, Ansible, Docker).
- Adaptability, passion to learn new areas and tools, and flexibility to work and communicate effectively across different teams and time zones.
Nice to have
- Practical experience with Infiniband/Ethernet networks in areas like RDMA, topologies, and congestion control.
- Experience debugging network issues in large-scale deployments.
- Familiarity with CUDA programming and/or GPUs.
- Experience with Deep Learning Frameworks such as PyTorch and TensorFlow.
Benefits
- Base salary range of 152,000 USD - 241,500 USD for Level 3.
- Base salary range of 184,000 USD - 287,500 USD for Level 4.
- Eligibility for equity and benefits.
Job details
- Applications accepted at least until October 7, 2026.
- Posting type: Existing vacancy.
About the company
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing, and Visualization.
The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services.
Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars.
Additional information
- Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.
- NVIDIA uses AI tools in its recruiting processes.
Application: UpVagas lists this opportunity from the hiring company or its recruitment platform. The application is completed at the original source.