Thorben Consulting, LLC is seeking a highly skilled Site Reliability Engineering (SRE) + Release Pipeline Engineer to take end-to-end ownership of the software delivery lifecycle. This role combines build, release, and operations responsibilities, ensuring reliability and performance in production environments while contributing to high-priority mission initiatives.
Responsibilities:
- Run an engineering practice driven by evaluation, testing, and verification — changes are proven with tests, traces, and metrics before they are called done
- Operate agentic systems fluently (MCP tools, Bedrock/LLM calls, streaming, distributed traces) in support of GenAI IDP document classification and field extraction workflows
- Work within DCSA classified environments (IL5/IL6) following DoD security baselines and STIG compliance requirements
- Support the GenAI IDP solution: prompt engineering, model fine-tuning, evaluation logic for document section classification and extracted field validation
- Use AI coding assistants heavily as daily drivers, while remaining skeptical of their output and holding it to the same evidentiary bar as any other code
- Developing STIG-compliant AMI automation, integrating with Artifactory, and enhancing GitLab CI-driven deployment capabilities
- Implementing baseline account pipelines, enforcing tagging compliance, and supporting centralized VPC endpoint configurations
- Enabling secure, repeatable Infrastructure-as-Code (IaC)-based deployments of the GenAI Intelligent Document Processing (IDP) solution across customer non-production and production environments
Requirements:
- MUST be a US Citizen
- MUST have an active DoD TOP SECRET clearance or be eligible to obtain one
- 3+ years of professional experience in SRE, DevOps, platform, or infrastructure engineering
- Experience operating Kubernetes (EKS) in production, including in DoD/IC classified environments
- Experience packaging and deploying applications with Helm (authoring and maintaining charts, not only consuming them)
- Experience with Flux (or an equivalent GitOps controller — e.g., Argo CD) driving continuous delivery of Helm releases
- Experience with AWS compute and managed services (e.g., EKS, RDS, S3, IAM/IRSA, EC2 Image Builder)
- Experience with infrastructure-as-code; TypeScript/CDK experience specifically, or demonstrated ability to work in a TypeScript IaC codebase
- Experience with production observability and distributed tracing (e.g., OpenTelemetry, Grafana/Tempo, or equivalent), used to diagnose failures from telemetry rather than guesswork
- Experience leading incident diagnosis and resolution, including identifying and confirming root cause before remediating
- Experience with GitLab CI/CD pipelines for build automation and deployment
- Proficiency in at least one scripting language (e.g., Python) for tooling and automation
- Current, active TS/SCI security clearance with polygraph
- Current, active Security+ or equivalent certificate for privileged user access
- Experience building STIG-compliant AMI pipelines using EC2 Image Builder with DoD security baseline validation
- Experience with Artifactory integration for AMI/container image distribution across AWS Organizations
- Experience with cross-domain solutions (AWS Diode, CDS) and SIPRNet/JWICS environments
- Experience with AWS Organizations, SCPs, OU design, and multi-account governance
- Experience with certificate lifecycle management (ACM, Private CA) and automated renewal workflows
- Experience deploying or operating ML/LLM-serving infrastructure (e.g., Amazon Bedrock endpoints, SageMaker, model-inference endpoints) and reasoning about token/latency/cost from traces
- Experience with an ECR- or registry-based GitOps model (charts/images mirrored to a registry, controller reconciling from it)
- Experience diagnosing failed or stuck Helm/Flux reconciliations (HelmRelease not converging, drift between desired and live state)
- Experience with Envoy/gateway and certificate/TLS management in classified environments
- Proficiency using AI coding assistants as a daily driver, with the judgment to validate their output before it reaches a cluster
- Bachelor's degree in Computer Science or a related field, or equivalent practical experience
- Bachelor's degree in Engineering, Computer Science, Information Systems, or related field
- Associate-level or higher AWS certification