01Key Responsibilities
Fleet Architecture & Technical Ownership
Own the endtoend technical architecture of hyperscale GPU fleets, including hardware platform selection, firmware strategy, OS configuration, drivers, networking, and observability.
Define and enforce technical standards and best practices for fleet reliability, availability, performance, and operability.
Lead major fleetwide initiatives such as new GPU platform bringups, multigeneration hardware transitions, and architectural redesigns.
Evaluate tradeoffs across cost, performance, reliability, and timetodeploy, and make technically sound decisions under ambiguity.
Reliability, Availability & Performance
Set and drive fleetlevel reliability, availability, and performance objectives.
Lead rootcause analysis and resolution of complex, systemic failures affecting large portions of the fleet or multiple datacenters.
Identify recurring failure patterns and drive longterm fixes spanning hardware, software, automation, and operational processes.
Work directly with hardware vendors and partners to resolve platformlevel issues and influence future hardware designs.
Automation & Systems Engineering
Design and build largescale automation systems for:
GPU fleet provisioning and lifecycle management
GPU health validation, diagnostics, and certification
Automated remediation, recovery, and replacement workflows
Eliminate manual operational .