Reddit, Inc. is a community of communities that fosters open and authentic conversations on the internet. They are seeking a Senior Machine Learning Engineer to lead optimization initiatives and improve efficiency for Ads ML workloads, collaborating closely with various teams to enhance model training and inference processes.
Responsibilities:
- Independently own high-value optimization initiatives across training, inference, or launch-readiness for important Ads ML workloads
- Diagnose bottlenecks in real production systems using profiling, benchmarking, and observability rather than intuition-first debugging
- Build performance tooling, optimization playbooks, observability hooks, guardrails, or efficiency primitives that help more than one team or workload over time
- Improve launch-safety and efficiency readiness by contributing to load testing, fallback readiness, latency and cost visibility, and operational confidence for heavy models
- Work with model owners and platform teams to land pragmatic fixes while helping the team gradually standardize repeated solutions
- Contribute to the team’s technical direction by surfacing patterns, tradeoffs, and opportunities for reuse or automation
- Mentor less-experienced engineers through code, debugging, measurement rigor, and strong execution habits
Requirements:
- Deep ML systems experience close to real production models and workloads, not just generic infra exposure
- Direct hands-on experience improving training or serving efficiency with measurable outcomes
- Strong technical judgment across model-level, runtime-level, and infrastructure-level optimization choices
- Ability to own complex projects end to end and collaborate effectively across team boundaries
- Good customer and platform instincts: can solve concrete bottlenecks while keeping maintainability, adoption, and future reuse in mind
- Strong communication: able to explain tradeoffs clearly to engineers and partner teams
- Experience with GPU training or serving migrations
- Experience with PyTorch, distributed training frameworks, or kernel/runtime optimization
- Experience building launch certification, efficiency benchmarking, or cost observability systems
- Experience in organizations where platform and applied modeling responsibilities are split across multiple teams
- Experience with model compression or deployment optimizations such as quantization, pruning, distillation, or checkpoint optimization