Back to jobs

Research Intern, Efficient Deep Learning – 2027

Internship US, CA, Santa Clara Hybrid Posted 2 horas ago

NVIDIA

Opportunity published from Workday · applications stay with the original hiring source.

Source verified 1 hour agoThe original job page was active during UpVagas' latest lifecycle check.

About this opportunity

About the opportunity: NVIDIA is seeking a PhD intern to join the Deep Learning Efficiency Research (DLER) team, focusing on efficient deep learning research that creates real-world impact.

The team concentrates on efficient diffusion language models, multimodal generative models, and efficient agentic AI with hybrid inference orchestration across cloud and edge, alongside post-training model optimization, efficient architecture design, adaptive inference, and resource-efficient training.

You will collaborate with a research team that regularly publishes at top computer vision and machine learning venues, contributing to projects that can directly influence NVIDIA products.

Key responsibilities

  • Research, design, and implement novel methods for efficient deep learning in diffusion LLMs, multimodal models, or efficient agentic AI
  • Develop sampling efficiency, adaptive unmasking, self-speculation, parallel decoding, and multimodal generation pipelines
  • Work on hybrid inference orchestration, routing and scheduling policies, expert delegation, and resource-aware agent loops
  • Publish original research and collaborate with internal team members, external researchers, and product groups to transfer technology

Qualifications

  • Pursuing a Ph.D. in Computer Science/Engineering, Electrical Engineering, or a related field
  • Excellent knowledge of the theory and practice of machine learning and deep learning
  • Experience with large language models, diffusion language models, multimodal or vision-language models, or agentic systems
  • Hands-on experience with large-scale model training including data preparation and model parallelization, such as tensor and pipeline parallelism
  • Outstanding research track record with at least one publication at a top-tier conference like ICML, ICLR, NeurIPS, CVPR, or ICCV
  • Excellent communication skills

Nice to have

  • Parallel programming experience, such as CUDA
  • Interest or experience in hybrid cloud-edge inference, orchestration, or adaptive routing
  • Background in pruning, quantization, NAS, or efficient backbones

Benefits

  • Eligible for standard NVIDIA intern benefits
  • Competitive hourly pay based on position, location, year in school, degree, and experience

Job details

  • Pay range: 38 USD - 94 USD per hour
  • Application deadline: Applications will be accepted at least until October 9, 2026

About the company: NVIDIA is widely considered one of the technology world's most desirable employers, featuring forward-thinking and hardworking teams experiencing rapid growth.

The company is driven by creative and autonomous innovators with a passion for computer architecture and advanced technology.

Additional information

  • This posting is for an existing vacancy.
  • NVIDIA uses AI tools in its recruiting processes.

Application: UpVagas lists this opportunity from the hiring company or its recruitment platform. The application is completed at the original source.

Source transparency. UpVagas organizes information provided by the company or recruitment source. Always review the original listing before submitting personal information.
Get new jobs on Telegram Tap the mascot to set your job alerts.
UpVagas mascot with Telegram