CoreFleet Solutions
US - California - San Jose
View Company Profile /
<< Go Back
****About the role****
* We are seeking an experienced RMA Failure Analysis for GPU Servers and enterprise server platforms. The engineer will be responsible for diagnosing, troubleshooting, and performing root cause analysis on customer-returned GPU servers, server motherboards, GPU baseboards, and associated hardware subsystems. This ideal candidate will possess strong server architecture knowledge, component-level debugging expertise, and the ability to safely handle and analyze high-value hardware throughout the failure analysis process.
****What you'll do****
* Perform failure analysis on customer-returned GPU servers, server motherboards, GPU boards, GPU baseboards, and related hardware assemblies.
* Conduct system-level, board-level, and component-level troubleshooting to identify root causes of hardware failures.
* Execute functional testing, diagnostics, and debug activities using standard lab equipment and server validation tools.
* Read and interpret schematics, block diagrams, board layouts, and manufacturing documentation.
* Analyze failures involving server subsystems including CPUs, GPUs, DIMMs, NICs, SSDs, power supplies, PCIe devices, and cooling/thermal subsystems.
* Troubleshoot hardware issues related to BIOS, BMC, CPLD, FPGA, PCIe, memory, storage, networking, and power delivery circuits.
* Perform component-level debugging including capacitors, resistors, fuses, diodes, MOSFETs, voltage regulators, ICs, and other electronic components.
* Conduct component swapping, isolation testing, and fault reproduction to validate failure mechanisms and root causes.
* Perform detailed visual and mechanical inspections to identify damaged, missing, misaligned, overheated, or improperly assembled components.
* Utilize JIRA and Zendesk to track RMA cases, document failure analysis results, manage issue resolution activities, and maintain clear communication across engineering, quality, and customer support teams.
* Document failure analysis findings, corrective actions, and recommendations to support continuous product quality improvements.
* Collaborate with design, validation, manufacturing, and quality teams to drive issue resolution and corrective actions.
* Follow proper ESD and hardware handling procedures while working with customer-returned products, engineering samples, and production hardware.
****Qualifications****
****Required Qualifications****
* 4 years of experience in server hardware design, validation, testing, debugging, failure analysis, or system engineering.
* Strong understanding of GPU server architecture and enterprise server platforms.
* Experience performing system-level, board-level, and component-level troubleshooting.
* Ability to read and interpret electrical schematics, block diagrams, and PCB layouts.
* Hands-on experience with server technologies including BIOS, BMC, CPLD, FPGA, PCIe, memory subsystems, storage interfaces, and networking interfaces.
* Experience using laboratory equipment such as oscilloscopes, digital multimeters (DMM), power analyzers, logic analyzers, and protocol analyzers.
* Working knowledge of Linux operating systems and command-line troubleshooting.
* Strong understanding of root cause analysis methodologies and failure isolation techniques.
* Ability to safely handle sensitive server and GPU hardware while adhering to ESD and hardware handling best practices.
****Preferred Qualifications****
* Experience supporting AI, HPC, or GPU-accelerated server platforms.
* Experience with customer-returned hardware (RMA) failure analysis processes.
* Knowledge of power delivery architecture, thermal analysis, and signal integrity concepts.
* Familiarity with manufacturing defects, field failures, and reliability-related investigations.
****Critical Requirements****
* Must be capable of independently troubleshooting GPU servers and server hardware down to the component level.
* Must understand overall server architecture and subsystem interactions before initiating debug activities.
* Must demonstrate strong analytical and problem-solving skills in hardware failure analysis.
* Must be comfortable working with customer-returned hardware and managing multiple RMA investigations simultaneously.
* Must maintain proper hardware handling practices to prevent damage to customer-returned units and engineering samples.
About CoreFleet Solutions
=========================
At CoreFleet Solutions, we're building a company focused on delivering exceptional workforce, logistics, and technology services. As a growing startup, every team member has the opportunity to make a meaningful impact and help shape the future of the business.
We partner with organizations to provide staffing solutions, logistics support, and technology deployment services with a commitment to quality, reliability, and customer success.
### Why Join CoreFleet?
* Opportunity to grow with a fast-growing startup
* Work directly with company leadership
* Learn new skills across multiple industries
* Collaborative, supportive, and entrepreneurial culture
* Make a real impact, your ideas and contributions matter
If you're looking for a place where you can grow your career while helping build something from the ground up, we'd love to hear from you.
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX