arbeit0815
DEEN
← Alle Stellen

Nvidia

Software Engineer Intern, AI and DL Kernel Libraries - 2027

China, Shanghai · Vor Ort · AI and DL Kernel Libraries · vor 21 Std.

NVIDIA is looking for outstanding Software Engineer Interns to help develop groundbreaking technologies for AI and deep learning kernel libraries. Our team builds core software that accelerates high-impact AI workloads on NVIDIA GPUs, with a strong focus on deep learning primitives, kernel libraries, and performance-critical GPU software. As an intern on the team, you will contribute to the design, development, optimization, and delivery of software that powers NVIDIA's AI platform.

This internship is centered on foundational library engineering, with opportunities to work on low-level kernels, performance primitives, and efficient implementations for modern AI and deep learning workloads. You may contribute to GPU-accelerated deep learning primitives, attention kernel implementations, runtime components, code generation systems, and other performance-critical infrastructure for large language models and advanced AI applications. You will collaborate with world-class engineers across deep learning software, compilers, GPU architecture, and open-source inference ecosystems, and your work can directly impact the performance of real-world workloads at scale.


What you'll be doing


  • Contribute to production-quality software that ships as part of NVIDIA's AI software stack, including cuDNN, FlashInfer, and optimized support for large language model inference workloads.
  • Help develop new AI systems technologies for efficient inference, with a focus on performance, scalability, maintainability, and usability.
  • Support the design, implementation, and optimization of kernels for high-impact AI workloads across LLM inference, generative AI, computer vision, autonomous driving, and recommender systems.
  • Assist in building extensible software abstractions for deep learning libraries, LLM serving engines, and runtime systems.
  • Contribute to just-in-time compilation, code generation, and runtime technologies for performance-critical GPU workloads.
  • Analyze workload performance, tune current software, and help propose improvements to future software and hardware-software interfaces.
  • Collaborate closely with engineers across deep learning frameworks, libraries, kernels, compilers, and GPU architecture teams at NVIDIA.
  • Contribute to open-source communities and ecosystem integrations where relevant, including projects such as FlashInfer, vLLM, and SGLang.

What we need to see


  • Currently pursuing a Bachelor's, Master's, or PhD degree in Computer Science, Electrical Engineering, or a related field.
  • Coursework, research, or hands-on project experience in machine learning, deep learning systems, compilers, systems software, or GPU programming.
  • Strong programming skills in C/C++ and Python.
  • Familiarity with CUDA development and GPU programming fundamentals.
  • Experience developing with or using deep learning frameworks such as PyTorch, JAX, TensorFlow, or ONNX.
  • Understanding of linear algebra, performance analysis, profiling, and code optimization.
  • Interest in software abstractions, APIs, and higher-level system architecture for performance-sensitive systems.
  • Interest in modern machine learning and inference system trends, especially around LLMs and generative AI.
  • Strong problem-solving skills, curiosity, and the ability to work effectively in a collaborative environment.

Ways to stand out from the crowd


  • Hands-on experience with inference engines and runtimes such as vLLM, SGLang, MLC, TensorRT-LLM, or similar systems.
  • Background in domain-specific compilers, code generation, or library solutions for LLM inference and training.
  • Exposure to machine learning compilers or IR systems such as MLIR, Apache TVM, TensorIR, or related technologies.
  • Practical experience with GPU performance modeling, computer architecture, or accelerator-oriented software design.
  • Open-source project ownership or meaningful contributions in deep learning systems, compilers, kernels, or inference infrastructure.

arbeit0815 · Daten werden alle 6 Stunden aktualisiert