RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference. This role involves working on core systems that optimize performance and cost across thousands of GPUs, directly shaping the deployment of state-of-the-art models for users worldwide.
Responsibilities:
- Design and build large-scale inference systems for frontier AI models
- Optimize latency, throughput, and GPU utilization in production inference
- Develop and improve model serving architectures and runtimes
- Work on batching, scheduling, and memory management strategies
- Collaborate with kernel, compiler, and systems teams on performance optimization
- Debug performance bottlenecks across the stack
- Drive reliability and scalability of inference infrastructure
- Build tooling for observability, profiling, and performance analysis
- Contribute to long-term inference architecture and strategy