Digital-first government for the common good.
DevOps Engineer IV, Operational Resilience, Observability, SRE
Location
United States
Posted
7 days ago
Salary
$120K - $150K / year
Seniority
Lead
Job Description
DevOps Engineer IV, Operational Resilience, Observability, SRE
Ad Hoc LLC
• DevOps Engineer IV serves as a senior individual contributor and technical leader within a team, providing leadership, guidance, and mentoring to other engineers • You will be responsible for meeting scope, schedule, and delivery requirements, interacting with stakeholders, and driving improvements in DevOps processes and practices across the program. • Set up and operate monitoring and observability tooling (AWS CloudWatch, Prometheus, Grafana, Loki log aggregation) for real-time visibility into application health, performance, and infrastructure • Build and maintain "Golden Signals" performance dashboards measuring latency, traffic, errors, and saturation • Implement the DORA metrics roadmap using Grafana and GitLab analytics to establish performance baselines • Manage Tier 2/3 production support within strict SLAs: 1-hour initial response, 4-hour critical resolution, 99.9% uptime commitment • Author Root Cause Analyses within 3 business days of any severity-1 production outage; maintain on-call runbooks and change correlation • Author and maintain the BCDR plan, including recovery architecture and RTO targets, cross-region replication (RDS, S3), Route 53 routing, and Secrets Manager; coordinate biannual failover drills • Configure centralized alerting and incident tooling (Jira Service Desk/ServiceNow, Microsoft Teams, AWS Chatbot) • Implement AWS Auto Scaling and Elastic Load Balancing; deliver sprint performance reports and cost-optimization recommendations • Support recruiting efforts by evaluating homework assignments and potentially assisting with interviews
Job Requirements
- Bachelor's degree and 8+ years of relevant experience, or equivalent additional experience in lieu of a degree
- Must meet federal suitability requirements and pass a background investigation as a condition of employment
- 5+ years of hands-on experience with AWS, Terraform (or similar IaC), and Git/GitLab in production environments
- Experience supporting 5 or more engineering teams from a shared DevOps/platform function
- Demonstrated experience designing and building CI/CD pipelines in GitLab and/or Jenkins, including quality and security gates
- Strong working knowledge of containerization using Docker and orchestration on AWS ECS/EKS/Fargate
- Familiarity with DevSecOps practices including SAST, dependency scanning, and automated vulnerability remediation
- Excellent communication and documentation skills; comfortable in a highly collaborative Agile/SAFe environment
- Proven experience in SRE, production operations, or incident response for mission-critical, high-availability cloud-native services
- BCDR planning and disaster recovery exercise experience
Benefits
- Company-subsidized health, dental, and vision insurance
- Flexible PTO
- 401K with employer match
- Paid parental leave after one year of service
- Employee Assistance Program
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
DevOps Engineer IV – Operational Resilience, Observability, SRE
Ad Hoc LLCDigital-first government for the common good.
• DevOps Engineer IV serves as a senior individual contributor and technical leader within a team, providing leadership, guidance, and mentoring to other engineers. • You will be responsible for meeting scope, schedule, and delivery requirements, interacting with stakeholders, and driving improvements in DevOps processes and practices across the program. • Set up and operate monitoring and observability tooling (AWS CloudWatch, Prometheus, Grafana, Loki log aggregation) for real-time visibility into application health, performance, and infrastructure • Build and maintain "Golden Signals" performance dashboards measuring latency, traffic, errors, and saturation • Implement the DORA metrics roadmap using Grafana and GitLab analytics to establish performance baselines • Manage Tier 2/3 production support within strict SLAs: 1-hour initial response, 4-hour critical resolution, 99.9% uptime commitment • Author Root Cause Analyses within 3 business days of any severity-1 production outage; maintain on-call runbooks and change correlation • Author and maintain the BCDR plan, including recovery architecture and RTO targets, cross-region replication (RDS, S3), Route 53 routing, and Secrets Manager; coordinate biannual failover drills • Configure centralized alerting and incident tooling (Jira Service Desk/ServiceNow, Microsoft Teams, AWS Chatbot) • Implement AWS Auto Scaling and Elastic Load Balancing; deliver sprint performance reports and cost-optimization recommendations • Support recruiting efforts by evaluating homework assignments and potentially assisting with interviews
DevOps Engineer IV – CI/CD Pipeline, Platform Engineering
Ad Hoc LLCDigital-first government for the common good.
• Design, build, and maintain automated CI/CD pipelines in GitLab CI and Jenkins with build-once promotion patterns, fail-fast verification, and automated rollbacks across dev, staging, and production • Create and maintain shared deployment tooling, reusable GitLab CI templates, and Terraform modules that 5+ AppDev delivery teams consume as self-service infrastructure • Maintain standard, secure base Docker images for Java, Python, and Angular applications; deploy containers on AWS ECS, EKS, or Fargate with zero-downtime patterns • Enforce pipeline security and quality gates (Trivy container scanning, SonarQube static analysis, Open Policy Agent policy-as-code) at the commit/merge stage • Mitigate over-privileged IAM configurations using least-privilege standards without disrupting active production services • Develop the roadmap for and implement DORA metrics (Deployment Frequency, Change Lead Time, Change Failure Rate, MTTR) in the program's CI/CD capabilities • Author and maintain runbooks, deployment documentation, and architecture documentation; lead knowledge-transfer sessions with engineering teams • Participate in SAFe ceremonies including PI planning, dependency management, and CI/CD roadmap development • Support recruiting efforts by evaluating homework assignments and potentially assisting with interviews
DevOps Engineer IV, CI/CD Pipeline, Platform Engineering
Ad Hoc LLCDigital-first government for the common good.
• DevOps Engineer IV serves as a senior individual contributor and technical leader within a team, providing leadership, guidance, and mentoring to other engineers. • You will be responsible for meeting scope, schedule, and delivery requirements, interacting with stakeholders, and driving improvements in DevOps processes and practices across the program. • Design, build, and maintain automated CI/CD pipelines in GitLab CI and Jenkins with build-once promotion patterns, fail-fast verification, and automated rollbacks across dev, staging, and production. • Create and maintain shared deployment tooling, reusable GitLab CI templates, and Terraform modules that 5+ AppDev delivery teams consume as self-service infrastructure. • Maintain standard, secure base Docker images for Java, Python, and Angular applications; deploy containers on AWS ECS, EKS, or Fargate with zero-downtime patterns. • Enforce pipeline security and quality gates (Trivy container scanning, SonarQube static analysis, Open Policy Agent policy-as-code) at the commit/merge stage. • Mitigate over-privileged IAM configurations using least-privilege standards without disrupting active production services. • Develop the roadmap for and implement DORA metrics (Deployment Frequency, Change Lead Time, Change Failure Rate, MTTR) in the program's CI/CD capabilities. • Author and maintain runbooks, deployment documentation, and architecture documentation; lead knowledge-transfer sessions with engineering teams. • Participate in SAFe ceremonies including PI planning, dependency management, and CI/CD roadmap development. • Support recruiting efforts by evaluating homework assignments and potentially assisting with interviews.
• function primarily as a Dev/Ops Administrator and Engineer for the Rocket Enterprise Server environment • interfacing with vendors for issue resolution and providing technical guidance to development teams • assist with developing infrastructure design specifications supporting critical legacy core applications • evaluate, plan, and integrate hardware, software, and middleware solutions while ensuring seamless system performance • apply and test patch updates and upgrades as required • install, configure, automate, and maintain Micro Focus Enterprise Server and Enterprise Developer across multiple environments on Linux • perform system maintenance including spool and log management, diagnostics, and issue resolution • design infrastructure solutions to meet business requirements for Micro Focus, Linux/Unix/Windows, and iSeries/AS400 environments • ensure IT governance compliance by managing technical documentation, KPI reporting, code promotions, third-party software updates, and disaster recovery requirements • optimize system performance by analyzing throughput and resource utilization, implementing necessary modifications • participate in on-call Tier 2 and 3 production support activities • coordinate activities with vendors as required

