Blue Sky Group is on a mission to transition the social web from platforms to protocols and is building a federated social network called AT Protocol. They are seeking a Staff Site Reliability Engineer to design, implement, and operate the infrastructure that powers their systems, ensuring reliability and operational excellence.
Responsibilities:
- You'll work across bare-metal systems, cloud services, data infrastructure, observability, incident response, capacity planning, and reliability engineering for systems serving millions of users
- Own reliability, availability, and operational excellence for our production systems, including observability, incident response, deployment, and rollback systems
- Improve production readiness for services, migrations, and infrastructure changes
- Develop software that pushes the state of the art in performance, automation, observability, and other areas
- Scale systems running on dense, latest-generation, bare-metal servers in our own colocation facilities
- Reduce toil through automation, tooling, and thoughtful engineering practices
- Partner with engineers across all our teams to help design services with strong operational characteristics
- Lead incident reviews and turn contributing factors into concrete engineering improvements as we practice continuous improvement
- Perform capacity planning and cost management across compute, storage, database, and networking workloads
- Manage various vendor relationships to ensure we can provide high quality services at a reasonable TCO
- Mentor engineers on reliability, operability, debugging, and distributed systems practices and help define a culture of operational excellence across the org
Requirements:
- Have +10 years experience operating high-scale production systems, including bare metal
- Have strong fundamentals in Linux, networking, storage, databases, and distributed systems
- Have built and operated high-scale systems where correctness, latency, throughput, and availability were critical
- Can write production-quality software in Go
- Are comfortable debugging across application code, operating systems, databases, networks, and hardware
- Have experience with observability systems, alert design, incident response, capacity planning, kubernetes, and production automation
- Like working on very small, fast-moving teams at a startup
- Have read the AT Protocol docs, feel aligned with the mission, and want to contribute!