Back to jobs

System Software Engineer, Distributed Systems

Full-time Remote - United States Remote Posted 2 horas ago

NVIDIA

Opportunity published from Workday · applications stay with the original hiring source.

Source verified 2 hours agoThe original job page was active during UpVagas' latest lifecycle check.

About this opportunity

About the opportunity: The VLSI Productivity and Infrastructure team at NVIDIA supports over 1,000 chip design engineers by building tools and platforms that supercharge their everyday work.

Our mission is to make chip designers faster through build automation, observability, analytics, automated error detection, and codebase modernization.

Our core workflow infrastructure runs as userspace software on bare-metal Linux hosts without sudo or containers, coordinating shared state and artifacts via NFS and launching compute-heavy workflows on IBM LSF.

We seek a pragmatic and versatile systems engineer with an emphasis on distributed systems and operational excellence to build tools that empower other engineers.

Key responsibilities

  • Design, build, and deliver core components of our next-generation productivity platforms
  • Develop reliable userspace infrastructure for long-running engineering workflows at scale on bare-metal Linux hosts
  • Build state coordination over NFS, focusing on atomicity, idempotency, dedup, and partial-write recovery without privileged ops
  • Build and improve orchestration around IBM LSF, including submission, tracking, retries, cancel, log capture, fairness, and backpressure
  • Convert legacy codebases into modern powerhouses using incremental migration techniques such as Perl to Go with stage gates and strong observability
  • Debug and improve performance and reliability across Linux and Kubernetes, including operational tooling
  • Collaborate with engineering users to turn ambiguous workflows into durable production systems

Qualifications

  • B.S. CS/EE or equivalent experience
  • 5+ years developing and operating production software in Go and/or Python, ideally in large codebases
  • Strong Linux fundamentals covering processes, filesystems, permissions, synchronization, locks, concurrency, and debugging
  • Solid distributed-systems thinking addressing failures, retries, timeouts, backoff, idempotency, and operational rigor
  • Experience building long-runtime automation or services on shared compute clusters, batch schedulers, or build systems
  • Ability to translate ambitious, high-level goals into a safe delivery plan with instrumentation, staged rollout, and measurable outcomes

Nice to have

  • Hands-on experience with shared filesystems at scale like NFS, or coordination patterns on eventually-consistent storage
  • Experience with batch job scheduling, shared compute fleets, or build systems
  • Track record of incremental modernization including tests, shadow runs, canaries, and rollback plans
  • Experience partitioning and optimizing metadata-heavy systems and reducing I/O or R/W hot spots
  • Strong incident and debug tactics including clear root-cause analysis, remediation, guardrails, and rapid comprehension of unfamiliar codebases

Benefits

  • Competitive salary range of 152,000 USD - 241,500 USD for Level 3 and 184,000 USD - 287,500 USD for Level 4
  • Equity eligibility
  • Comprehensive benefits package

Job details

  • Base salary determined by location, experience, and peer pay
  • Applications accepted at least until October 10, 2026

About the company: NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years through great technology and amazing people.

Today, we are tapping into the unlimited potential of AI to define the next era of computing where our GPU acts as the brains of computers, robots, and self-driving cars.

As an NVIDIAN, you will be immersed in a diverse, supportive environment where everyone is inspired to do their best work.

Additional information

  • NVIDIA uses AI tools in its recruiting processes

Application: UpVagas lists this opportunity from the hiring company or its recruitment platform. The application is completed at the original source.

Source transparency. UpVagas organizes information provided by the company or recruitment source. Always review the original listing before submitting personal information.
Get new jobs on Telegram Tap the mascot to set your job alerts.
UpVagas mascot with Telegram