Nvidia
Manager, Infrastructure Engineering and DevOps
Israel, Yokneam · Onsite · Infrastructure Engineering and DevOps · 8h ago
NVIDIA is looking for an outstanding Manager, Infrastructure Engineering - Networking to lead a high-impact infrastructure engineering team. The team develops and maintains engineering infrastructure solutions that enable R&D teams to provision, validate, test, debug, and recover complex server and networking environments at scale.
In this role, you will lead a team that provides infrastructure platforms and hands-on engineering support for internal customers across firmware, driver, hardware, software, and verification organizations. The position combines deep technical leadership with people management, execution ownership, customer support, and operational excellence in a fast-paced R&D environment. We are looking for a strong technical manager who can grow and mentor engineers, set technical direction, drive recovery of complex systems, optimize customer flows, and deliver reliable infrastructure capabilities at NVIDIA scale.
What you’ll be doing:
Lead and grow an infrastructure engineering team responsible for bare-metal provisioning, VM infrastructure, server fleet automation, CI/CD infrastructure, customer-facing debug support, and high-performance networking environments.
Own the team’s technical roadmap, priorities, execution plans, and delivery commitments across multiple infrastructure initiatives, while balancing long-term platform improvements with day-to-day customer needs.
Drive ownership of VM box inventory and lifecycle management across many Linux distributions, including image readiness, OS compatibility, package baselines, kernel configurations, provisioning flows, and production availability.
Build infrastructure capabilities that enable engineering and verification teams to run provisioning, testing, validation, and debug workflows efficiently and reliably, without positioning the team as the owner of verification itself.
Lead customer support, debug, and optimization of internal customer flows, including triage, root-cause analysis, bottleneck removal, workflow improvements, and clear communication with engineering stakeholders.
Guide complex system debug and recovery in a firmware R&D environment, including server bring-up issues, driver and firmware interactions, boot failures, networking problems, lab instability, automation failures, and environment recovery.
Provide technical leadership for Linux-based automation platforms, including server lifecycle management, OS installation, kernel configuration, driver setup, inventory management, resource allocation, observability, and production readiness.
Partner closely with firmware, driver, hardware, software, cloud, and verification teams to define requirements, improve reliability, and deliver infrastructure solutions that accelerate engineering productivity.
What we need to see:
B.Sc. in Computer Engineering, Computer Science, Electrical Engineering, or a related technical field, or equivalent experience.
8+ overall years of experience in Linux systems administration, infrastructure automation, DevOps, system software, firmware infrastructure, lab infrastructure, or related engineering domains.
3+ years of experience leading or managing engineering teams, technical projects, or cross-functional infrastructure initiatives.
Strong technical background in Linux environments, including systemd, package management, kernel parameters, GRUB, sysctl tuning, NFS, networking, boot flows, and service management.
Hands-on experience designing, implementing, and debugging automation software using Python, scripting, CI/CD workflows, and modern software development practices.
Experience managing infrastructure across multiple Linux distributions, including OS image management, compatibility issues, provisioning flows, package dependencies, and environment consistency.
Proven ability to support internal customers in complex technical environments, including issue triage, root-cause analysis, flow optimization, incident handling, and communication with cross-functional R&D teams.
Strong people leadership skills, including coaching, mentoring, performance management, hiring, feedback, prioritization, and building an inclusive, high-performing team.
Ways to stand out from the crowd:
Experience managing teams that build infrastructure platforms for firmware R&D, hardware bring-up, driver development, lab automation, cloud provisioning, or large-scale engineering environments.
Deep knowledge of high-speed networking technologies such as RDMA, InfiniBand, Ethernet, OFED, SR-IOV, PCI passthrough, VFIO/IOMMU, or related Linux networking tools.
Experience with Ansible, infrastructure-as-code practices, Jenkins, Kubernetes, Docker, KVM, QEMU, libvirt, Vagrant, or multi-architecture environments such as x86_64, aarch64, and ppc64le.
Familiarity with NVIDIA/Mellanox hardware, including ConnectX NICs, BlueField DPUs, firmware tools, hardware diagnostics, RSHIM, MFT, BIOS/BMC automation, or server recovery workflows.
Experience with complex system recovery and fleet operations, including Redfish, iDRAC, iLO, IPMI, BIOS automation, BMC configuration, remote power control, automated remediation, and production incident reduction.
With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us, and our engineering teams are growing rapidly. If you are a technical leader with a passion for infrastructure, automation, Linux systems, networking, customer enablement, and building strong engineering teams, we want to hear from you.
NVIDIA is committed to fostering a diverse work environment and is proud to be an equal opportunity employer. We highly value diversity in our current and future employees and do not discriminate, including in our hiring and promotion practices, on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status, or any other characteristic protected by law.