arbeit0815
DEEN
← Alle Stellen

Nvidia

Senior Data Infrastructure Engineer, AI Performance

China, Shanghai · Vor Ort · AI Computing Architecture · vor 16 Std.

NVIDIA’s AI Computing Architecture team develops analytical models and simulators that guide the design of future GPUs, systems, and AI platforms. These tools enable architects to explore large design spaces and understand performance, power, cost, and software trade-offs across constantly evolving AI workloads.
 

We are seeking outstanding software engineers to build and scale the infrastructure behind this simulation ecosystem. You will develop the distributed execution, data, automation, and visualization platforms that turn architectural models into reliable, reproducible, large-scale studies. You will be sitting at the intersection of distributed systems, performance engineering, data platforms, and architecture.
 

What you’ll be doing:

  • Build scalable and reliable infrastructure for running large simulation studies across on-premises compute clusters and cloud environments, improving throughput, resource efficiency, and reproducibility.

  • Establish a unified storage and data platform as the source of truth for simulation configurations, execution state, results, and provenance.

  • Develop self-service analytics and visualization capabilities that help architects explore results and compare performance, power, and design trade-offs.

  • Partner with GPU architects, performance engineers, AI researchers, and software teams to translate emerging AI workloads into reusable simulation platform capabilities.

  • Help define the technical roadmap and engineering practices for NVIDIA’s next generation of AI architecture simulation platforms.
     

What we need to see:

  • BS or higher degree in a relevant technical field (CS, EE, CE, Math, etc.).

  • 3+ years of experience building production infrastructure, distributed systems, or data platforms.

  • Strong software engineering and system-design skills in one or more programming languages, with a solid understanding of scalability, reliability, data consistency, and operational trade-offs.

  • Ability to work through ambiguous problems, simplify fragmented systems, and collaborate effectively across engineering and research teams.
     

Ways to stand out from the crowd:

  • Deep expertise in distributed systems, storage and query engines, compute orchestration, or large-scale data platforms.

  • Experience improving the performance, reliability, or efficiency of compute intensive and data intensive systems.

  • Understanding of LLM inference optimization and the performance characteristics of conversational, agentic, or other emerging AI workloads.

  • Familiarity with architecture simulation, performance modeling, high-performance computing, GPU computing, or AI workloads.

  • A record of strong technical ownership and impact through industry work, research or open-source contributions.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

arbeit0815 · Daten werden alle 6 Stunden aktualisiert