Noble Machines builds multipurpose robots to support human workers in challenging jobs. They are seeking an experienced ML Ops & Infrastructure Engineer to develop the foundational systems for AI development, focusing on scalable compute, data platforms, and deployment pipelines.
Responsibilities:
- Design, build, and maintain a highly scalable and reliable machine learning infrastructure that accelerates the research and development lifecycle
- Architect and manage robust data ingestion, collection, and processing pipelines. You will own the data platforms that ensure our models are trained on high-quality, perfectly versioned datasets
- Build and optimize the environments used for distributed model training, hyperparameter tuning, and automated model evaluation
- Manage and orchestrate heavy compute workflows seamlessly across AWS and/or Google Cloud Platform (GCP), optimizing for both performance and cost
- Take full ownership of containerizing ML workloads and orchestrating them via Kubernetes (K8s) to ensure high availability, scalability, and reproducibility
- Partner closely with ML Researchers and Software Engineers to understand their bottlenecks, gather requirements, and build tooling that makes their workflows frictionless