NEAR AI is building decentralized and confidential machine learning infrastructure to enable user-owned AI. They are specifically seeking an expert in high-performance LLM serving systems and inference optimization.
Responsibilities:
Architect and maintain production high-traffic LLM serving systems
Optimize throughput, latency, and cost for leading open-source LLMs