Software for rapid military planning: make planning fast enough for today's environment
Senior Site Reliability Engineer (Arlington, VA) - Relocation Provided
Location
United States
Posted
5 days ago
Salary
$180K - $220K / year
Seniority
Senior
Job Description
Senior Site Reliability Engineer (Arlington, VA) - Relocation Provided
Onebrief
Consequential Work. Dedicated People. About Onebrief Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination. Today, many critical planning workflows still rely on fragmented systems, static documents, and disconnected tools that make collaboration and decision-making unnecessarily difficult. Onebrief brings modern software, AI, and real-time collaboration into those environments, helping teams operate with greater clarity, coordination, and adaptability in situations where decisions carry real-world consequences. We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world. Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth. Security Clearance, Location, and Onsite Notice:This role requires regularly working on-site at customer locations in Arlington, VA. If you are not currently within commuting distance, you must be willing to relocate (note that Onebrief will provide relocation assistance). Active Top Secret Clearance required with the ability to obtain SCI eligibility. About The RoleWe are hiring a Site Reliability Engineer to join our Infrastructure & Security team. You’ll work closely with fellow SREs, security, and customer success. You will be the first line of support for our mission critical deployments, and responsible for ensuring best-in-class service quality and issue resolution. You will work in both on-premise DoD environments and AWS cloud environments. Your lessons from the field will shape how our team works, from policy to implementation. In addition to working at the customer, you will contribute directly to solutions that increase stability, performance, and security of our deployments, and improve the overall experience of deploying and managing Onebrief on premise. About YouYou care deeply about reliability and treat it as a core feature of any application or platform, with a bias toward “reliability over novelty.” You think about infrastructure and operability as products to be automated, well-documented, and continuously improved, and you aim to leave systems easier to operate than you found them. You are equally comfortable leading a post-incident review, or diving into a kubectl shell to triage a complex production issue. You don't just fix problems; you translate constraints and failure modes into clear, automated guardrails and scalable, resilient architecture. For you, robust monitoring, actionable alerting, and insightful runbooks are core parts of the engineering process, not afterthoughts. You mentor others, fostering a culture of blameless postmortems and proactive reliability. You collaborate naturally with application and platform teams, helping them move quickly but safely by building the tools, processes, and observability that make "fast recovery" a reality. What You'll DoYou'll own the reliability, scalability, and security of the production application and/or platform. You will do this by: - Implementing a World-Class Observability Platform: Design, implement, and manage our monitoring, logging, and alerting stack (e.g., Prometheus, Loki, Alloy, and Grafana). You won't just track metrics; you'll create the actionable insights and automated alerting that allow teams to identify and resolve issues before they impact users. - Defining and Upholding Reliability: Define, measure, and own alerting that feeds into our Service Level Indicators (SLIs) and Service Level Objectives (SLOs), increasing trust internally and externally. You will be the organization's expert on what it means for our systems to be reliable and how to measure it. - Leading Incident Response: Act as the incident responder and potentially incident commander during critical incidents who will lead blameless post-mortems / After Action Reviews (AARs) that identify true root causes and drive automated, long-term solutions to prevent recurrence. - Automating for Scale and Security: Partner with platform engineers to design, build, and manage secure, resilient Kubernetes clusters and cloud/on-prem environments using Infrastructure-as-Code (Terraform, Ansible). You will embed security and compliance controls (RMF, STIGs) directly into this automation. - Eliminating Toil and Scaling the Team: Proactively identify and eliminate operational toil by building automation. You will partner with other teams to share best practices for air-gapped environments and support their readiness for production. What We Look For - An active Top Secret clearance - 5+ years in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus. - Proven partner to DevOps/Platform and application teams; collaborates well across functions and shares context openly. - A deep understanding of incident response processes, with experience conducting thorough root cause analyses and driving continuous improvement. Technical expertise - Infrastructure as Code: Terraform (or CloudFormation), Ansible. - Containers and orchestration: Kubernetes design, deployment, and operations. - CI/CD: experience building and maintaining pipelines (GitLab CI/CD, Jenkins, GitHub Actions). - Scripting: proficiency with at least one of Python, Go, or Bash. - Cloud: Familiarity with AWS or AWS GovCloud. - Observability: Grafana stack, ELK stack, or Datadog. - Networking fundamentals: core protocols and secure configurations. Bonus points (nice to have) - Experience in DoD environments and compliance frameworks (RMF, STIGs, ICD 503). - GitOps practices and toolchains. - Security‑minded design for sensitive environments. - Experience designing and implementing meaningful SLIs/SLOs (including error budgets) for complex, distributed systems. - Familiarity with on‑prem virtualization(VMware, Proxmox, Nutanix, Hyper-V, etc). - Service mesh exposure (Istio, Linkerd). - Relevant certifications (e.g., AWS DevOps Engineer, CKA/CKAD). - Active Security+ or another DoD 8570.01-approved security credential, or the ability to obtain the valid credentials within 3 months of employment. Notice to Third Party Recruitment Agencies Please note that Onebrief does not accept unsolicited resumes from recruiters or employment agencies. In the absence of an executed Recruitment Services Agreement, there will be no obligation to any referral compensation or recruiter fee. In the event a recruiter or agency submits a resume or candidate without an agreement Onebrief explicitly reserves the right to pursue and hire those candidate(s) without any financial obligation to the recruiter or agency. Any unsolicited resumes, including those submitted to hiring managers, shall be deemed the property of Onebrief.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
GCP DevOps Engineer
Ruri Software Technologies LLCMAGIC serves as the centralized ERP system for over 120 State agencies, supporting: Financial Management Procurement Grants Management Reporting These modules have been in production since 2015. The State is also currently implementing SAP SuccessFactors for: Human Capital Management (HCM) Payroll Processing
Role Description We are seeking a GCP DevOps Engineer for a 12+ month project with VZ. The ideal candidate will have extensive experience in building pipelines and I/CD (Continuous Integration and Continuous Delivery). - Experience in build automation tools like Jenkins, Docker, Kubernetes. - Expert in using different source code version control tools like Git. - Installation and configuration of DNS, Load Balancer, SSL, HTTP, FTP, and TCP/IP. - Experience with containers and orchestration services like Kubernetes and Docker. - Experience in public cloud architecture and engineering in Google Cloud Platform (GCP). - Experience in configuration, deployment, management, and maintenance of large cloud-hosted systems, including scaling, monitoring, performance tuning, troubleshooting, and disaster recovery. - Experience with Linux/UNIX environments and scripting for build & release automation. - Application deployments and environment configuration using Terraform and Ansible. - Involved in installing, configuring, and administering Hadoop clusters. - Hadoop administration, Spark tuning, data rebalancing, and tuning MapReduce and Spark jobs. - Extensively worked with Dataproc, GCS, Dataflow, GKE, BigQuery, JupyterLab, Jupyter Notebook, Vertex AI, and Domino. Qualifications - Proven experience in DevOps practices and methodologies. - Strong understanding of cloud architecture and services. - Familiarity with containerization and orchestration tools. Requirements - Minimum of 3 years of experience in a DevOps role. - Strong knowledge of Google Cloud Platform. - Experience with automation and configuration management tools. Benefits - Competitive salary. - Flexible working hours. - Opportunity for professional development.
Role Description We are seeking a skilled and experienced Site Observability Engineer to join the NS2 Observability team. The ideal candidate will be responsible for improving our monitoring and alerting posture for Cloud Infrastructure. The role requires a strong understanding of observability tools and practices, with a focus on: - Prometheus - Grafana - Gardener Kubernetes - Splunk Experience with Dynatrace is a plus. - Implement, manage, and improve monitoring solutions that use Prometheus, ensuring high availability and accurate alerting for our systems. - Contribute to the development of observability strategies to improve our Cloud monitoring posture. - Collaborate with development teams to integrate observability into the CI/CD pipeline and throughout the application lifecycle. - Respond to and investigate incidents, providing thorough post-mortem analyses and implementing preventive measures. - Stay current with the latest trends and best practices in site reliability and observability. - Work with cross-functional teams to ensure system reliability, scalability, and performance. Qualifications - Proven experience with observability tools such as Prometheus, Grafana, and Splunk. - Hands-on experience with Kubernetes and container orchestration, preferably with Gardener Kubernetes. - Familiarity with logging, monitoring, and application performance management (APM) tools; experience with Dynatrace is a plus. - Strong understanding of cloud infrastructure, networking, and distributed systems. - Excellent problem-solving and analytical skills, with the ability to work independently and as part of a team. - Strong communication skills and the ability to work effectively with both technical and non-technical stakeholders. - Experience with scripting and automation tools (Python, Terraform, Ansible, etc.). - Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent experience.
Role Description Cloud DevOps SME responsible for ensuring secure, repeatable, auditable deployments using Infrastructure-as-Code and delivery automation. - Evaluation Emphasis: - Terraform or equivalent IaC depth - Secure delivery implementation in federal environments - Integration of SAST/SCA/container scanning - Policy-as-code and configuration baseline management - Prevention of configuration drift - Core Responsibilities: - Design IaC templates for onboarding - Integrate delivery automation pipelines - Enforce policy-as-code controls - Maintain version-controlled repositories - Support continuous ATO monitoring - Ensure environment parity Qualifications - Minimum 7 years Software Delivery and Security Operations experience - Strong IaC expertise - Kubernetes/OpenShift experience - Experience in regulated environments Requirements - IRS MBI Clearance is required Company Description
Senior DevOps Engineer – Platform Engineering
Allied Technology ServicesThis is an exciting opportunity to work on modern cloud security initiatives, protect enterprise-level infrastructure, and collaborate with global teams in a fast-paced and security-focused environment.
Role Description Are you passionate about building scalable platforms, automating software delivery, and improving developer experience? Join us as a Senior DevOps Engineer – Platform Engineering and play a key role in modernizing the software delivery lifecycle for our mission-critical Election Management platforms. This is a highly technical, hands-on engineering role where you'll design and build modern DevOps solutions—not just maintain existing infrastructure. If you enjoy creating CI/CD pipelines, developing automation, implementing Infrastructure as Code, and building engineering platforms that improve reliability and productivity, we'd love to hear from you. - Design and modernize enterprise CI/CD pipelines using Azure DevOps. - Build deployment automation and reusable engineering frameworks. - Implement Infrastructure as Code (Terraform, CloudFormation, or similar). - Improve cloud architecture, scalability, resiliency, and operational efficiency on AWS. - Develop automation using Python, PowerShell, Bash, Go, or similar languages. - Implement observability solutions (monitoring, logging, tracing, dashboards, and alerting). - Support Kubernetes and containerized workloads. - Drive DevSecOps, security automation, and disaster recovery initiatives. - Improve developer experience through self-service platform capabilities. - Collaborate with Engineering, QA, Security, Product, and Architecture teams to establish modern DevOps and Platform Engineering best practices. Qualifications - 8+ years of experience in DevOps, Platform Engineering, Site Reliability Engineering (SRE), or Software Engineering. - Strong hands-on experience with Azure DevOps and AWS. - Expertise in CI/CD, Infrastructure as Code, and cloud-native architectures. - Experience with Kubernetes and containerized applications. - Strong scripting/programming skills (Python, PowerShell, Bash, Go, or similar). - Experience implementing monitoring, observability, and deployment automation. - Excellent problem-solving, communication, and stakeholder management skills. Requirements - Bonus Points - Experience with GitOps (ArgoCD). - Knowledge of Datadog, Grafana, CloudWatch, ELK, or OpenTelemetry. - Experience with DevSecOps and security automation. - Background in AI-assisted operations (AIOps). - Experience with FedRAMP, GovCloud, or regulated cloud environments. - Previous work supporting government, election, or mission-critical platforms. Benefits At CIVIX, you'll help build secure, scalable technology that supports critical public services. You'll work alongside talented engineers, influence platform strategy, and have the opportunity to shape the future of our engineering practices while solving complex technical challenges.

