Thinking Machines Lab is dedicated to advancing collaborative general intelligence and creating accessible AI tools. They are seeking a Reliability Engineer to ensure the reliability of their GPU supercomputing fleet, focusing on diagnosing hardware issues and collaborating with vendors to maintain operational efficiency.
Responsibilities:
- Investigate, reproduce, and remediate issues across large GPU clusters
- Own the drivers, kernel surface, and diagnostics that span hardware, firmware, and OS
- Automate the monitoring of fleet reliability and analyze error rates to validate whether a fix or firmware change measurably reduced failures rather than shifting them around
- Drive the firmware lifecycle: tracking, qualification, staged rollout, and regression analysis
- Engage vendors directly — GPUs, server OEMs, NIC vendors, and storage vendors — to get real fixes rather than ticket numbers. Manage RMA flows when hardware needs to come out
- Monitor and improve GPU hardware health signals and turn them into actionable reliability improvements
- Write clear postmortems and vendor cases that move issues forward