Discogs is the largest crowd-sourced, community-driven database of recorded music information in the world. They are seeking a Senior Site Reliability Engineer to contribute to the Platform team’s centralized infrastructure, focusing on maintenance, monitoring, and automation of services, while also mentoring engineering squads and driving improvements in technologies and processes.
Responsibilities:
- Owning tasks and larger projects from planning to production rollout
- Learning new technologies and building expertise with the goal of teaching and mentoring others; mentoring with the goal of force-multiplying through docs and tools
- Maintaining organization cloud presence in AWS
- Automating and deploying infrastructure configurations using Infrastructure as Code (IAC)
- Mentoring engineering squads on Platform best practices for Kubernetes, MySQL, Kafka, and other software development lifecycle areas
- Assisting engineering squads with capacity planning, on-call preparation, and production readiness
- Writing documentation and runbooks that contribute to the engineering organization’s knowledge base
- Implementing monitoring and alerting systems with Discogs observability tools
- Working in a containerized, orchestrated environment
- Participating in on-call rotation, responding to incidents, and troubleshooting data and other operations issues
- Contributing to the reliability and design patterns of our Kafka CDC and event workflows
- Contributing to agentic AI best practices and tooling, including skills, agents, and safety
Requirements:
- Infrastructure-as-code (Terraform)
- CI/CD (GitHub Actions)
- Kubernetes (EKS, Kustomize, Karpenter, administration, application manifests)
- AWS and cloud development (VPC, EKS, RDS, S3)
- FinOps and cloud cost optimization
- Observability (Datadog, Sentry)
- Agentic AI (Claude Code)
- Scripting (Shell, Python)
- Track record of collaboration and mentorship
- Excellent written communication and documentation skills
- Continuous learning
- Ownership and proactive approach to solving large problems
- A Bachelor's Degree in Computer Science or similar area of focus, or equivalent relevant work experience
- 5+ years experience in Ops, DevOps, Site Reliability, Platform or other systems roles
- Kafka: Cluster administration (Strimzi), Kafka Connect (Debezium, JDBC)
- Flink
- Relational database administration and performance (MySQL, Percona Server, AWS RDS)
- Elasticsearch (ECK administration, scaling, performance)
- Python (SQLAlchemy, FastAPI)
- GraphQL (schema design, Apollo federation)
- REST API
- GitOps (ArgoCD)
- Hashicorp Vault
- Redis
- Memcached