MyFitnessPal is a company dedicated to enabling users to make healthy choices through their innovative tools and resources. They are seeking a Software Engineer III - Site Reliability to enhance the reliability and security of their software delivery systems while collaborating within the PEAS team to ensure a seamless user experience.
Responsibilities:
- Own and evolve our SLI/SLO and error-budget frameworks, and use them to influence prioritization and product decisions
- Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches
- Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue
- Design and operate resilient, scalable infrastructure using Infrastructure as Code (Terraform)
- Manage production Kubernetes and container workloads, including capacity planning and cloud-cost optimization
- Own CI/CD pipelines and safe deployment strategies (canary, progressive rollout, fast rollback)
- Own the security controls that live inside the delivery pipeline — integrating and tuning SAST, DAST, and SCA scanning (for example, in GitHub Actions) so issues surface while code is still in review
- Implement and maintain policy-as-code (for example, OPA/Rego, Kyverno, or Conftest) to block unsafe infrastructure and Kubernetes changes at admission time
- Drive vulnerability triage and remediation SLAs for pipeline- and infrastructure-level findings, prioritizing by real risk
- Partner with our Security Engineer and the broader Security & Reliability disciplines — you own security in the pipeline and collaborate on the rest, rather than duplicating that function
- Participate in and improve the on-call rotation; build the runbooks and automation that make on-call sustainable
- Coach team members and engineers across the org on reliability patterns and operational best practices
Requirements:
- 5+ years in site reliability, platform, or infrastructure engineering, with clear senior-level ownership of production systems
- Strong programming skills for automation and tooling (Go, Python, Typescript or similar) — you have experience building software or custom tooling, not just scripts
- Deep, hands-on experience with a major cloud platform (AWS is a plus), Kubernetes, and Infrastructure as Code (Terraform is a plus)
- Proven track record leading incident response and building SLO-driven reliability practices
- Working fluency with observability tooling (Datadog is a plus)
- Practical experience integrating security into CI/CD pipelines — SAST/DAST/SCA tooling, dependency scanning, or policy-as-code
- Strong understanding of cloud security fundamentals (identity/IAM, least-privilege patterns, policy/guardrails, secrets management)
- The judgment and communication skills to raise a security or reliability finding with a senior engineer and land it as a shared problem to solve, not a fight to win
- Experience with policy-as-code frameworks (especially Kyverno, but tools like OPA/Rego or Conftest are also relevant) enforced at admission time is a plus
- Exposure to regulated or compliance-driven environments (SOC 2, PCI DSS, HIPAA) is a plus
- Chaos engineering or game-day experience is a plus
- Experience supporting B2C/mobile backend environments with high traffic, rapid iteration, and strong reliability needs is a plus