xAI logo

xAI

Hardware Failure Analysis Engineer - Memphis at xAI

Southaven, MS; Memphis, TNFull-timeData CenterPosted about 1 month ago
Apply with Pipeline

About the Role

<div class="content-intro"><p><span style="font-family: arial, helvetica, sans-serif;">SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge.&nbsp;</span><span style="font-family: arial, helvetica, sans-serif;">Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. </span><span style="font-family: arial, helvetica, sans-serif;">We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. </span><span style="font-family: arial, helvetica, sans-serif;">All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.</span></p></div><h3 data-pm-slice="1 1 []"><span style="font-family: arial, helvetica, sans-serif;">ABOUT THE ROLE:</span></h3> <p>Determine true root cause of fleet hardware failures — component and system level — and drive fixes through vendors to the manufacturer. Turn "swap it again" into "vendor redesign / firmware fix / manufacturing escape found."</p> <h3><span style="font-family: arial, helvetica, sans-serif;">RESPONSIBILITIES:</span></h3> <ul> <li>Own named failure classes across GPU trays/baseboards, NIC/DPU, motherboard/PCIe switch, memory, power, thermal, cables/connectors, and rack-scale patterns.</li> <li>Run recurrence analysis and fleet-wide defect clustering; detect systemic patterns before they become fleet-scale loss.</li> <li>Build vendor technical escalation packages with evidence quality that forces action; partner with OEM/ODM/component manufacturers on firmware, bring-up, and manufacturing escapes through committed CAPA.</li> <li>Feed findings into RMA policy, spare strategy, and "do not reseat forever" stop-rules.</li> <li>Work the FA intake queue on rotation; keep queue age within SLA</li> </ul> <h3><span style="font-family: arial, helvetica, sans-serif;">BASIC QUALIFICATIONS:</span></h3> <ul> <li>Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field (or equivalent experience).</li> <li>2+ years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments.</li> <li>Proven expertise in firmware analysis, hardware specifications review, and release validation.</li> <li>Strong experience with RMA processes, including filing claims, vendor negotiations, and pushing for resolutions outside standard protocols.</li> <li>Demonstrated ability to diagnose and prove complex hardware failures, including grey or intermittent issues, using tools, logic analyzers, or diagnostic software.</li> <li>Familiarity with data center hardware components (e.g., servers, GPUs, networking equipment) and emerging technologies.</li> <li>Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Rust, or similar). Not required to be expert in all of them.</li> <li>Excellent problem-solving skills with a data-driven approach to reliability engineering.</li> <li>Ability to work collaboratively with cross-functional teams, including operations technicians.</li> </ul> <h3><span style="font-family: arial, helvetica, sans-serif;">PREFERRED SKILLS AND EXPERIENCE:</span></h3> <ul> <li>Experience in AI/ML infrastructure or supercomputing environments.</li> <li>Knowledge of vendor ecosystems (e.g., NVIDIA, Dell, HP, Supermicro) and supply chain management.</li> <li>Certifications in hardware engineering or reliability (e.g., CRE, CompTIA Server+).</li> <li>Prior work in a fast-paced startup or tech company like SpaceXAI.</li> </ul><div class="content-conclusion"><p><em>SpaceXAI is an equal opportunity employer. For details on data processing, view our </em><em><a href="https://x.ai/legal/recruitment-privacy-notice" target="_blank">Recruitment Privacy Notice</a>.</em></p></div>