Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence. They are seeking an infrastructure research engineer to design and build core systems for efficient large-scale model training, focusing on numerics and distributed training.
Responsibilities:
- Design and optimize distributed training infrastructure for large-scale LLMs, focusing on performance, stability, and reproducibility across multi-GPU and multi-node setups
- Implement and evaluate low-precision numerics (for example, BF16, MXFP8, NVFP4) to improve efficiency without sacrificing model quality
- Develop kernels and communication primitives that use hardware-level support for mixed and low-precision arithmetic
- Collaborate with research teams to co-design model architectures and training recipes that align with emerging numeric formats and stability constraints
- Prototype and benchmark scaling strategies such as data, tensor, and pipeline parallelism that integrate precision-adaptive computation and quantized communication
- Contribute to the design of our internal orchestration and monitoring systems to ensure that thousands of distributed experiments can run efficiently and reproducibly
- Publish and share learnings through internal documentation, open-source libraries, or technical reports that advance the field of scalable AI infrastructure