Back to jobs

Senior Software Engineer – Distributed Systems Engineer, EDA Infrastructure

Full-time Remote - United States Remote Posted 2 horas ago

NVIDIA

Opportunity published from Workday · applications stay with the original hiring source.

Source verified 1 hour agoThe original job page was active during UpVagas' latest lifecycle check.

About this opportunity

Sr Software Engineer - Distributed Systems Engineer, EDA Infrastructure

NVIDIA is hiring engineers to build and scale the infrastructure that supports our Electronic Design Automation (EDA) workloads.

We are looking for engineers with strong programming skills, a deep understanding of distributed systems, experience operating large-scale production infrastructure, and excellent communication and planning abilities.

You will help design reliable automation and platform services that manage large fleets of GPU-based and CPU-based compute systems used by engineering teams across NVIDIA.

What You Will Be Doing

  • Design and build platforms that automate the provisioning, configuration, operation, and lifecycle management of large-scale GPU and CPU compute infrastructure.
  • Develop monitoring, health-management, and remediation systems that improve the reliability, availability, and utilization of EDA compute environments.
  • Automate hardware deployment, operating-system configuration, firmware and software updates, cluster enrollment, and recovery workflows.
  • Build reliable services and workflows that integrate with workload schedulers, infrastructure management systems, and observability platforms.
  • Use hardware diagnostics, operating-system signals, scheduler data, and network and storage telemetry to identify failures and return unhealthy systems to service.
  • Work with EDA, infrastructure, networking, storage, and hardware engineering teams to deliver scalable solutions for critical chip-design workloads.
  • Participate in incident response, root-cause analysis, capacity planning, and the continuous improvement of production services.

What We Need To See

5+ years of software engineering or infrastructure engineering experience supporting large-scale production systems.

A BS in Computer Science, Engineering, Physics, Mathematics, or a related field, or equivalent experience.

Strong programming experience in Go or Python, including a solid understanding of data structures, algorithms, testing, and software design.

Experience designing automation for distributed systems and large fleets of Linux-based compute nodes.

Understanding of performance, security, reliability, fault tolerance, state management, and data consistency in complex systems.

Ways To Stand Out From The Crowd

Experience designing or operating large-scale EDA or high-performance computing infrastructure. Deep knowledge of Linux, GPU and CPU server architecture, networking, storage, and bare-metal lifecycle management.

Hands-on experience with workload schedulers and cluster-management platforms such as Slurm, LSF, Kubernetes, or Bright Cluster Manager.

Experience supporting EDA applications, license-management systems, high-throughput batch workloads, or semiconductor design workflows.

Experience building automated health checks, break-fix remediation, firmware and operating-system upgrade workflows, or node-provisioning systems.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

Key details from the source

  • The ideal candidate is comfortable working across software, operating systems, cluster schedulers, networking, storage, and physical hardware.
  • Experience with infrastructure automation, software deployment, observability, and operational recovery.
  • You will also be eligible for equity and benefits .

Full listing

UpVagas condensed repetitive or legal boilerplate to make this page easier to read. The original hiring source contains the complete job description.

Source transparency. UpVagas organizes information provided by the company or recruitment source. Always review the original listing before submitting personal information.
Get new jobs on Telegram Tap the mascot to set your job alerts.
UpVagas mascot with Telegram