LearningSpring is The School Choice Management Platform that helps various stakeholders run modern education freedom programs. They are seeking a Site Reliability Engineer who will enhance systems and automate tasks, focusing on AWS-based platform operations and internal technology management.
Responsibilities:
- Administer and maintain internal business systems including Google Workspace, Slack, GitHub, Linear, Jira, Atlassian, Zoom, Outline, Customer.io, HubSpot, and other SaaS platforms
- Own identity and access management, including user provisioning, SSO, MFA, access reviews, and employee onboarding and offboarding
- Automate repetitive operational and administrative tasks using Bash, Python, APIs, and workflow automation
- Assist with security and compliance initiatives, including SOC 2 activities, Drata, access controls, endpoint security, and audit preparation
- Support and improve our AWS platform, including ECS, RDS, Route 53, and related cloud infrastructure
- Contribute infrastructure improvements using Terraform and GitHub Actions under the guidance of senior platform engineers
- Build, maintain, and improve monitoring, alerting, dashboards, and operational tooling using AWS-native services and future third-party observability platforms as needed
- Participate in incident response, troubleshooting, root cause analysis, and continuous operational improvement
- Improve deployment reliability, CI/CD processes, and engineering workflows
- Help optimize cloud infrastructure performance and cost
- Create and maintain operational documentation, runbooks, and internal knowledge resources
- Collaborate across Engineering and Operations to continuously improve reliability, security, and employee productivity
Requirements:
- 2–5 years of experience in Site Reliability Engineering, Platform Engineering, Systems Administration, DevOps, or a similar operations-focused role
- Proven working knowledge of Google Workspace
- Experience working with AWS cloud services in a production environment
- Experience with Docker and containerized applications
- Familiarity with Infrastructure as Code concepts, preferably Terraform
- Experience using GitHub and GitHub Actions
- Strong understanding of Linux system administration fundamentals, networking, DNS, TLS certificates, and cloud infrastructure concepts
- Experience administering cloud-based productivity platforms such as Google Workspace and modern SaaS applications
- Strong understanding of identity and access management, SSO, and MFA
- Experience writing automation scripts using Bash, Python, or similar languages
- Excellent troubleshooting, documentation, and communication skills
- A collaborative, low-ego mindset with a willingness to learn, take ownership, and contribute wherever needed
- Experience supporting SOC 2 or similar security and compliance frameworks
- Experience supporting AWS ECS-based environments
- Experience with AWS monitoring and observability tools such as CloudWatch
- Familiarity with PagerDuty or similar incident management platforms
- Experience integrating SaaS applications using APIs or workflow automation
- Startup or high-growth technology company experience
- Interest in growing into a mid-level platform or Site Reliability Engineering role