Ad Hoc LLC logo
Ad Hoc LLC

Digital-first government for the common good.

DevOps Engineer IV – Operational Resilience, Observability, SRE

DevOps EngineerDevOps EngineerFull TimeRemoteLeadTeam 501-1,000Since 2014H1B No SponsorCompany SiteLinkedIn

Location

Virginia

Posted

7 days ago

Salary

$120K - $150K / year

Seniority

Lead

Job Description

DevOps Engineer IV – Operational Resilience, Observability, SRE

Ad Hoc LLC

• DevOps Engineer IV serves as a senior individual contributor and technical leader within a team, providing leadership, guidance, and mentoring to other engineers. • You will be responsible for meeting scope, schedule, and delivery requirements, interacting with stakeholders, and driving improvements in DevOps processes and practices across the program. • Set up and operate monitoring and observability tooling (AWS CloudWatch, Prometheus, Grafana, Loki log aggregation) for real-time visibility into application health, performance, and infrastructure • Build and maintain "Golden Signals" performance dashboards measuring latency, traffic, errors, and saturation • Implement the DORA metrics roadmap using Grafana and GitLab analytics to establish performance baselines • Manage Tier 2/3 production support within strict SLAs: 1-hour initial response, 4-hour critical resolution, 99.9% uptime commitment • Author Root Cause Analyses within 3 business days of any severity-1 production outage; maintain on-call runbooks and change correlation • Author and maintain the BCDR plan, including recovery architecture and RTO targets, cross-region replication (RDS, S3), Route 53 routing, and Secrets Manager; coordinate biannual failover drills • Configure centralized alerting and incident tooling (Jira Service Desk/ServiceNow, Microsoft Teams, AWS Chatbot) • Implement AWS Auto Scaling and Elastic Load Balancing; deliver sprint performance reports and cost-optimization recommendations • Support recruiting efforts by evaluating homework assignments and potentially assisting with interviews

Job Requirements

  • Bachelor's degree and 8+ years of relevant experience, or equivalent additional experience in lieu of a degree
  • Must meet federal suitability requirements and pass a background investigation as a condition of employment
  • 5+ years of hands-on experience with AWS, Terraform (or similar IaC), and Git/GitLab in production environments
  • Experience supporting 5 or more engineering teams from a shared DevOps/platform function
  • Demonstrated experience designing and building CI/CD pipelines in GitLab and/or Jenkins, including quality and security gates
  • Strong working knowledge of containerization using Docker and orchestration on AWS ECS/EKS/Fargate
  • Familiarity with DevSecOps practices including SAST, dependency scanning, and automated vulnerability remediation
  • Excellent communication and documentation skills; comfortable in a highly collaborative Agile/SAFe environment
  • Proven experience in SRE, production operations, or incident response for mission-critical, high-availability cloud-native services
  • BCDR planning and disaster recovery exercise experience

Benefits

  • Company-subsidized health, dental, and vision insurance
  • Flexible PTO
  • 401K with employer match
  • Paid parental leave after one year of service
  • Employee Assistance Program

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Full TimeRemoteTeam 501-1,000Since 2014H1B No Sponsor

• Design, build, and maintain automated CI/CD pipelines in GitLab CI and Jenkins with build-once promotion patterns, fail-fast verification, and automated rollbacks across dev, staging, and production • Create and maintain shared deployment tooling, reusable GitLab CI templates, and Terraform modules that 5+ AppDev delivery teams consume as self-service infrastructure • Maintain standard, secure base Docker images for Java, Python, and Angular applications; deploy containers on AWS ECS, EKS, or Fargate with zero-downtime patterns • Enforce pipeline security and quality gates (Trivy container scanning, SonarQube static analysis, Open Policy Agent policy-as-code) at the commit/merge stage • Mitigate over-privileged IAM configurations using least-privilege standards without disrupting active production services • Develop the roadmap for and implement DORA metrics (Deployment Frequency, Change Lead Time, Change Failure Rate, MTTR) in the program's CI/CD capabilities • Author and maintain runbooks, deployment documentation, and architecture documentation; lead knowledge-transfer sessions with engineering teams • Participate in SAFe ceremonies including PI planning, dependency management, and CI/CD roadmap development • Support recruiting efforts by evaluating homework assignments and potentially assisting with interviews

United States
$120K - $150K / year
Ad Hoc LLC logo

DevOps Engineer IV, CI/CD Pipeline, Platform Engineering

Ad Hoc LLC

Digital-first government for the common good.

DevOps Engineer7 days ago
Full TimeRemoteTeam 501-1,000Since 2014H1B No Sponsor

• DevOps Engineer IV serves as a senior individual contributor and technical leader within a team, providing leadership, guidance, and mentoring to other engineers. • You will be responsible for meeting scope, schedule, and delivery requirements, interacting with stakeholders, and driving improvements in DevOps processes and practices across the program. • Design, build, and maintain automated CI/CD pipelines in GitLab CI and Jenkins with build-once promotion patterns, fail-fast verification, and automated rollbacks across dev, staging, and production. • Create and maintain shared deployment tooling, reusable GitLab CI templates, and Terraform modules that 5+ AppDev delivery teams consume as self-service infrastructure. • Maintain standard, secure base Docker images for Java, Python, and Angular applications; deploy containers on AWS ECS, EKS, or Fargate with zero-downtime patterns. • Enforce pipeline security and quality gates (Trivy container scanning, SonarQube static analysis, Open Policy Agent policy-as-code) at the commit/merge stage. • Mitigate over-privileged IAM configurations using least-privilege standards without disrupting active production services. • Develop the roadmap for and implement DORA metrics (Deployment Frequency, Change Lead Time, Change Failure Rate, MTTR) in the program's CI/CD capabilities. • Author and maintain runbooks, deployment documentation, and architecture documentation; lead knowledge-transfer sessions with engineering teams. • Participate in SAFe ceremonies including PI planning, dependency management, and CI/CD roadmap development. • Support recruiting efforts by evaluating homework assignments and potentially assisting with interviews.

Virginia
$120K - $150K / year
Full TimeRemoteTeam 10,001+H1B No Sponsor

• function primarily as a Dev/Ops Administrator and Engineer for the Rocket Enterprise Server environment • interfacing with vendors for issue resolution and providing technical guidance to development teams • assist with developing infrastructure design specifications supporting critical legacy core applications • evaluate, plan, and integrate hardware, software, and middleware solutions while ensuring seamless system performance • apply and test patch updates and upgrades as required • install, configure, automate, and maintain Micro Focus Enterprise Server and Enterprise Developer across multiple environments on Linux • perform system maintenance including spool and log management, diagnostics, and issue resolution • design infrastructure solutions to meet business requirements for Micro Focus, Linux/Unix/Windows, and iSeries/AS400 environments • ensure IT governance compliance by managing technical documentation, KPI reporting, code promotions, third-party software updates, and disaster recovery requirements • optimize system performance by analyzing throughput and resource utilization, implementing necessary modifications • participate in on-call Tier 2 and 3 production support activities • coordinate activities with vendors as required

Illinois
$78.7K - $125.9K / year
HostPapa logo

Site Reliability Engineer

HostPapa

Let Papa take care of you!

DevOps Engineer7 days ago
Full TimeRemoteTeam 51-200Since 2006H1B No Sponsor

• Define and implement SLIs, SLOs, and error budgets for critical CloudBlue services to ensure reliability and performance • Influence system architecture with a strong focus on reliability, scalability, and operability, designing systems for fault tolerance, graceful degradation, and self-healing • Reduce operational toil by identifying opportunities for automation and process improvement • Design and operate CloudBlue’s observability stack across metrics, logs, and traces using tools such as Datadog, Grafana, and Elastic Stack • Develop actionable alerting strategies and dashboards that provide clear insight into platform and business health • Design and maintain high-availability architectures, implementing redundancy, failover, and disaster recovery strategies across regions and availability zones • Conduct capacity planning, load testing, and performance optimization to ensure platform stability and scalability • Act as a senior responder during production incidents, leading incident coordination, communication, and service restoration • Own blameless postmortems and drive improvements that reduce incident frequency, MTTR, and customer impact • Improve reliability of Kubernetes-based platforms through health checks, autoscaling strategies, rollout safety, and resilience testing • Partner with engineering and DevOps teams to improve deployment safety, rollback strategies, and platform reliability • Maintain runbooks and operational documentation, and promote SRE best practices across engineering teams • Support other tasks or projects as assigned to meet team and business needs

Spain