- Home
- Jobs
- Data Center
- Hardware Failure Analysis Engineer - Memphis

Hardware Failure Analysis Engineer - Memphis at xAI
Southaven, MS; Memphis, TNFull-timeData CenterPosted about 1 month ago
Apply with PipelineAbout the Role
<div class="content-intro"><p><span style="font-family: arial, helvetica, sans-serif;">SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. </span><span style="font-family: arial, helvetica, sans-serif;">Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. </span><span style="font-family: arial, helvetica, sans-serif;">We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. </span><span style="font-family: arial, helvetica, sans-serif;">All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.</span></p></div><h3 data-pm-slice="1 1 []"><span style="font-family: arial, helvetica, sans-serif;">ABOUT THE ROLE:</span></h3>
<p>Determine true root cause of fleet hardware failures — component and system level — and drive fixes through vendors to the manufacturer. Turn "swap it again" into "vendor redesign / firmware fix / manufacturing escape found."</p>
<h3><span style="font-family: arial, helvetica, sans-serif;">RESPONSIBILITIES:</span></h3>
<ul>
<li>Own named failure classes across GPU trays/baseboards, NIC/DPU, motherboard/PCIe switch, memory, power, thermal, cables/connectors, and rack-scale patterns.</li>
<li>Run recurrence analysis and fleet-wide defect clustering; detect systemic patterns before they become fleet-scale loss.</li>
<li>Build vendor technical escalation packages with evidence quality that forces action; partner with OEM/ODM/component manufacturers on firmware, bring-up, and manufacturing escapes through committed CAPA.</li>
<li>Feed findings into RMA policy, spare strategy, and "do not reseat forever" stop-rules.</li>
<li>Work the FA intake queue on rotation; keep queue age within SLA</li>
</ul>
<h3><span style="font-family: arial, helvetica, sans-serif;">BASIC QUALIFICATIONS:</span></h3>
<ul>
<li>Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field (or equivalent experience).</li>
<li>2+ years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments.</li>
<li>Proven expertise in firmware analysis, hardware specifications review, and release validation.</li>
<li>Strong experience with RMA processes, including filing claims, vendor negotiations, and pushing for resolutions outside standard protocols.</li>
<li>Demonstrated ability to diagnose and prove complex hardware failures, including grey or intermittent issues, using tools, logic analyzers, or diagnostic software.</li>
<li>Familiarity with data center hardware components (e.g., servers, GPUs, networking equipment) and emerging technologies.</li>
<li>Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Rust, or similar). Not required to be expert in all of them.</li>
<li>Excellent problem-solving skills with a data-driven approach to reliability engineering.</li>
<li>Ability to work collaboratively with cross-functional teams, including operations technicians.</li>
</ul>
<h3><span style="font-family: arial, helvetica, sans-serif;">PREFERRED SKILLS AND EXPERIENCE:</span></h3>
<ul>
<li>Experience in AI/ML infrastructure or supercomputing environments.</li>
<li>Knowledge of vendor ecosystems (e.g., NVIDIA, Dell, HP, Supermicro) and supply chain management.</li>
<li>Certifications in hardware engineering or reliability (e.g., CRE, CompTIA Server+).</li>
<li>Prior work in a fast-paced startup or tech company like SpaceXAI.</li>
</ul><div class="content-conclusion"><p><em>SpaceXAI is an equal opportunity employer. For details on data processing, view our </em><em><a href="https://x.ai/legal/recruitment-privacy-notice" target="_blank">Recruitment Privacy Notice</a>.</em></p></div>
Related Roles
Data Center Manager
xAI
Southaven, MS; Memphis, TNData Center Operations Supervisor
xAI
Memphis, TNDriver (CDL) - Memphis (Night Shift)
xAI
Southaven, MS; Memphis, TNCritical Facilities Engineer - Memphis/Southaven
xAI
Southaven, MS; Memphis, TNLead Critical Facilities Technician - Memphis/Southaven
xAI
Southaven, MS; Memphis, TNSenior Critical Facilities Technician - Memphis
xAI
Southaven, MS; Memphis, TN