River AI is on a mission to create personal AI owned and shaped by each individual. They are seeking exceptional systems engineers to build high-performance engines that train their models, focusing on making training fast, reliable, and massively scalable.
Responsibilities:
- Architect and deploy fault-tolerant distributed systems for training and inference workloads across clusters with thousands of nodes
- Design high-performance kernels to maximize tensor operation efficiency, memory throughput, and networking over InfiniBand/RDMA
- Profile systems end-to-end to resolve blockers across hardware, software, data loading pipelines, and collective communication primitives
- Partner directly with research scientists to rapidly implement, optimize, and scale experimental model architectures