Nvidia
Senior System Management Architect
Israel, Yokneam · Vor Ort · Host Management Team · vor 11 Std.
Our Host Management Team focuses on AI factory management infrastructure. We define the next-generation management architecture of the data center — top-down, clusters, racks, platforms and multi-SoC trays. The architecture covers end-to-end flows for power-on, reset and boot sequences, power saving, management protocols to keep the system operating securely and reliably.
We are seeking a Senior System Management Architect to join our growing team. You will be the resident expert for system management, ensuring that AI factory management components (BMCs, MCUs, CPUs, network switches and CPLDs) operate in perfect unison. You will own the lifecycle management specifications for AI factory platforms — from factory provisioning to day-one deployment. If you enjoy studying a high-level system diagram and then immediately diving into the system-level buses and interface signals topology and the associated host management protocols modeling required to make it work, this role is for you.
What you'll be doing:
System Architecture & Flow Definition: Define and document comprehensive production, provisioning, update, recovery, and reset flows for AI factory sub-systems.
Ensure system-level architectural alignment and orchestrate seamless coordination between all the AI-Factory building blocks: GPU, DPU, CPU, Switch ASIC, NIC ASIC, BMC, and peripheral components.Protocol Implementation & Guidance: Specify the use of standard management protocols (Red-Fish, NSM, PLDM, SPDM, MCTP, NC-SI) across sub-system interfaces to implement the management tasks.
Hardware/Software Intersect: Define system-level hardware and software behavior, including CPLD logic requirements, power-sequencing dependencies, and reset orchestration (e.g., handling PCIe PERST#, SBR, and forced resets).
Interface Management: Architect intra-board communication pathways using PCIe, I2C/I3C, USB, and SPI, ensuring robust out-of-band (OOB) and in-band management connectivity.
Cross-Functional Leadership: Collaborate with Production, Hardware and Firmware Engineering, ASIC FW/Driver teams, SW/NOS teams and QA to translate architectural specifications into implementable, testable solutions.
Technical Documentation: Author rigorous technical specifications, test plans, and architectural design documents to guide development and factory production.
What we need to see:
Deep System-Level Perspective: Proven ability to understand complex computing or networking systems top-to-bottom, with a background in networking architectures and efficiency and power saving methods.
Protocol Fluency: Strong working knowledge of modern platform management and security protocols.
Hardware Interface Knowledge: Familiarity with board-level communication buses (I2C, I3C, SPI, USB) and system interconnects (PCIe architecture, including link states and reset mechanisms).
Analytical Documentation Skills: Exceptional ability to write clear, unambiguous technical specifications, state-machine descriptions, and sequence diagrams.
Experience: BS/MS in Electrical Engineering, Computer Engineering, Computer Science, or a related field, with 5+ years of relevant industry experience.
Ways to stand out from the crowd:
Experience with power-saving methods.
Background with Redfish RESTful API definitions and OOB network infrastructure.
Familiarity with hardware RoT concepts and secure boot architectures.
Experience with production line tooling, factory provisioning, and manufacturing test flows.
Background in networking principles, Switch, NIC/SmartNIC data-plane and offload operations, or high-speed data center topologies.