onXmaps, Inc. is a high-growth tech company focused on helping people confidently explore the outdoors. They are seeking a Site Reliability Engineer to build and maintain infrastructure that supports reliable software deployment at scale, ensuring systems are performant and accessible for development teams.
Responsibilities:
- Deploy, monitor and maintain highly available systems using technologies such as Terraform, CockroachDB and GCP services to include GKE(Kubernetes), Cloud SQL, Bigtable, Google Composer (Airflow), Google Cloud Storage, BigQuery, Pub/Sub, Cloud Run, etc
- Maintain and extend a large, mature Terraform codebase
- Analyze systems and make recommendations to increase performance, availability and minimize cost
- Automate manual systems to minimize toil wherever possible
- Develop and maintain integrations with 3rd party monitoring and alerting systems, such as Google Cloud Monitoring, Prometheus, OpenTelemetry, Checkly, and Rootly
- Drive incident response best practices for on-call engineering teams across onX
- Participate in the SRE team's on-call rotation for core infrastructure
- Collaborate in architectural decisions and direction involving our services and initiatives
Requirements:
- You have a B.S. or M.S. in computer science or a related field or relevant experience
- You have at least 5+ years of experience where 3+ are supporting production systems
- You have a strong interest and experience with Kubernetes, networking, and infrastructure-as-code
- You have experience with Terraform/OpenTofu
- You have exposure to at least one major cloud platform
- You evaluate technologies and solutions based on merit, stability, performance and the ability to debug
- You have practical experience with different types of datastores (SQL, NoSQL, object storage) and can explain when to use each based on data access patterns and scalability needs
- You have a strong computer science foundation
- You believe that your profession is a craft and you're driven to improve every day
- You take strong ownership of your work and platform responsibilities
- Familiarity with Google Cloud Platform
- Strong ability to troubleshoot and break down issues
- Experience working with high throughput, low latency services
- Experience working with a distributed team
- Experience working with IAM, auditing & security management within a cloud environment
- Experience working with GIS Mapping systems and tiles
- Experience working with Claude Code
- Experience working with Airflow or equivalent ETL systems