DevOps Engineer Remote Jobs in Alaska (US)
This page tracks remote devops engineer openings that are location-eligible for Alaska.
This page tracks remote devops engineer openings that are location-eligible for Alaska.
Open jobs
2,143
Hiring companies this week
10
Salary sample
$150,000 - $180,000
Jobs added last hour
0
2143 Jobs
1340 Companies
Software for rapid military planning: make planning fast enough for today's environment
Consequential Work. Dedicated People. About Onebrief Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination. Today, many critical planning workflows still rely on fragmented systems, static documents, and disconnected tools that make collaboration and decision-making unnecessarily difficult. Onebrief brings modern software, AI, and real-time collaboration into those environments, helping teams operate with greater clarity, coordination, and adaptability in situations where decisions carry real-world consequences. We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world. Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth. Security Clearance, Location, and Onsite Notice:This role requires regularly working on-site at customer locations in Arlington, VA. If you are not currently within commuting distance, you must be willing to relocate (note that Onebrief will provide relocation assistance). Active Secret Clearance required. About The RoleWe're hiring a Site Reliability Engineer to join our Infrastructure & Security team. You'll work closely with product engineers, fellow SREs, security, and customer success. This is an SRE role for someone who's comfortable in application code. Much of the reliability and performance work happens in the codebase (primarily TypeScript), so you'll fix problems at the source rather than working around them in the infrastructure. You'll be a first line of support for our mission-critical deployments across on-prem DoD and AWS environments, and what you learn in the field will feed directly back into the product. You'll ship code that makes Onebrief more stable, faster, and easier to deploy and operate. The work sits at the seam between engineering and operations, and it's weighted toward engineering. About YouYou treat reliability as a feature, not an afterthought, and you'd rather fix a problem in the code than route around it. You understand the full software development lifecycle (design, review, testing, release) and you know where reliability fits into each step. You're comfortable reading and writing application code, and you're just as happy dropping into a kubectl shell to triage a production issue. You turn failure modes into guardrails, and you think monitoring, alerting, and clear runbooks are part of building software, not extra credit. You mentor others and push a culture of blameless postmortems. You work naturally with product and platform teams, helping them move fast without breaking things by giving them the tools, tests, and observability that make quick recovery real. What You'll DoYou'll help make our production application reliable, scalable, and secure by improving the software itself, not just the systems it runs on. Day to day that looks like: - Improving the application: Work directly in the codebase (primarily TypeScript) to fix reliability and performance problems at the source. You'll partner with product engineers on design decisions, review code with reliability and security in mind, and treat "make the app better" as a first-class part of the job rather than something you hand off. - Building observability that developers actually use: Design and run our monitoring, logging, and alerting (Prometheus, Loki, Alloy, Grafana). The goal is alerts and dashboards tied to real application behavior, so teams catch issues before users do. - Owning reliability targets: Define and measure SLIs and SLOs, wire up alerting that feeds them, and be the person who can say what "reliable" means for our systems and prove it with data. - Leading incident response: Act as incident responder, and incident commander when needed. Run blameless post-mortems (AARs) that find the actual root cause and turn it into a code or process fix so it doesn't happen again. - Automating away toil: Spot the repetitive operational work and write software to kill it. Share what works with other teams, including those running in air-gapped environments, and help them get production-ready. What We Look For - An active Secret clearance - 5+ years in software engineering, SRE, or a related role, with real time spent writing and shipping application code - Strong TypeScript (or comparable modern language experience with willingness to work primarily in TypeScript) - Solid grasp of the full SDLC: design, code review, testing, release, and how reliability fits into each stage - Experience with incident response, root cause analysis, and turning findings into lasting fixes - A collaborator who works well across product, platform, and DevOps teams and shares context openly Technical expertise - Application development in TypeScript (Node and/or a modern front-end framework) - CI/CD: building and maintaining pipelines (GitHub Actions, GitLab CI/CD, Jenkins) - Testing and quality practices as part of the delivery process - Comfort with at least one of Python, Go, or Bash for tooling and automation - Working knowledge of containers and Kubernetes (enough to debug and deploy, not necessarily to stand up clusters from scratch) - Networking fundamentals and secure configuration basics Bonus points (nice to have) - Observability: Grafana stack, ELK, or Datadog - Infrastructure as Code (Terraform, Ansible) and cloud experience (AWS or AWS GovCloud) - Kubernetes cluster design and operations - Designing meaningful SLIs/SLOs with error budgets for distributed systems - GitOps practices and toolchains - DoD environments and compliance frameworks (RMF, STIGs, ICD 503) - Service mesh (Istio, Linkerd) - On-prem virtualization (VMware, Proxmox, Nutanix, Hyper-V) - Relevant certs (AWS DevOps Engineer, CKA/CKAD) Notice to Third Party Recruitment Agencies Please note that Onebrief does not accept unsolicited resumes from recruiters or employment agencies. In the absence of an executed Recruitment Services Agreement, there will be no obligation to any referral compensation or recruiter fee. In the event a recruiter or agency submits a resume or candidate without an agreement Onebrief explicitly reserves the right to pursue and hire those candidate(s) without any financial obligation to the recruiter or agency. Any unsolicited resumes, including those submitted to hiring managers, shall be deemed the property of Onebrief.
The Operating Core for Legal
• Own platform reliability, availability, and performance • Build AWS infrastructure using Terraform and Infrastructure as Code • Develop CI/CD pipelines, monitoring, and automation • Define SLOs, SLIs, and error budgets • Lead incident response and continuous reliability improvements • Mentor engineers and participate in a 24/7 on-call rotation
We help government agencies deliver public services as modern digital products.
• Lead the design, governance, and evolution of enterprise cloud architecture within a federal environment • Define cloud architecture standards and guide DevOps and engineering teams • Ensure Azure environments are designed for scalability, resilience, security, and compliance • Collaborate closely with engineering teams and government stakeholders to align cloud solutions with mission objectives • Design and oversee highly available, secure, and scalable Azure architectures • Define enterprise Azure reference architectures, patterns, and best practices • Establish architectural guidance for Azure identity, governance, policy enforcement, and monitoring frameworks • Evaluate and guide the adoption of cloud-native Azure services and modernization strategies • Define architectural standards for Infrastructure as Code (IaC) implementations • Guide DevOps teams in implementing automated CI/CD pipelines • Ensure IaC approaches support repeatable, secure, and auditable infrastructure deployment patterns • Provide architectural oversight of container-based deployment models • Lead architecture planning for migration of on-premises workloads to Azure • Develop strategies for hybrid connectivity and networking architectures • Define Azure governance strategies and ensure architecture aligns with federal security and compliance requirements • Serve as the technical authority for Azure architecture across the program
• Own CI/CD pipelines and release engineering across all environments (Dev/QA/Staging/Production/Demo/Multi-Region) • Own cloud infrastructure-as-code (reliability, scaling, and cost optimization on AWS) • Own the observability/monitoring stack, on-call alerting, and incident response • Implement and operate change management for production deployments • Own the information security program, policies, and the SOC 2 control set and evidence • Govern identity and access management (IAM) - provisioning standards, least privilege, and MFA enforcement • Run quarterly access reviews/recertification across all in-scope systems • Own vendor/third-party security risk assessment and review • Manage the SOC 2 auditor relationship and readiness milestones • Manage AWS Cognito identity integration, backups, and disaster-recovery procedures • Monitor platform uptime percentage (target 99.9%) • Own deployment frequency and change-failure rate; mean time to recovery (MTTR) • Manage infrastructure cost as a percentage of ARR • SOC 2 Type 1 readiness milestones delivered on schedule • Quarterly access reviews completed on time, exceptions remediated • Mean time to revoke access on offboarding (target = 1 business day) • Security incident count and time-to-contain
eDiscovery | Digital Forensics | Managed Review | AI & Analytics | Information Governance
• Develop and maintain Terraform, Ansible, and other infrastructure as code as required for automated infrastructure deployments, software deployments, and configuration management • Integrate deployments with existing management, security, and monitoring solutions • Develop webhooks between various systems leveraging APIs to achieve the desired integrations • Develop unit, functional and integration tests within code and CI/CD pipeline deployments to ensure automation code is resulting in the desired outcome • Integrate security controls and vulnerability testing into written code, automated CI/CD pipelines deployments, API integrations, and automated workflows using available tools • Review any security findings and work through remediation of findings • Leverage available orchestration tools to orchestrate the deployment of entire applications or service offerings • Troubleshoot errors and follow established policies within the orchestrated workflows • Develop and maintain Python, PowerShell or bash scripts for system integrations, system management or report generation of cloud infrastructure • Create and maintain comprehensive documentation such as diagrams, Wiki documentation, and code documentation • Mentor junior staff and provide training to enhance their skills in automation and script development • Adapt to various assigned duties in evolving technology scenarios.
• Drive the evolution of the company’s DevSecOps function by defining security initiatives, long-term priorities, and promoting a security-first engineering mindset. • Partner with the DevOps team to improve the security and resilience of Kubernetes (EKS), Istio, Service Mesh, and related cloud-native infrastructure. • Support engineering teams during service transitions by embedding security best practices into operational processes and ensuring effective knowledge sharing. • Develop and maintain policy-as-code frameworks and infrastructure validation mechanisms using technologies such as OPA and other compliance automation tools. • Improve the security of infrastructure provisioning and deployment workflows across Terraform, Ansible, and other Infrastructure-as-Code solutions. • Integrate, maintain, and optimise security testing and scanning capabilities throughout the software development and release lifecycle. • Build automation around security monitoring, compliance checks, and operational controls to increase efficiency and reduce manual effort.
Great infrastructure automation starts with great infrastructure data.
• Partner directly with customers and with our Solution Architecture/Customer Success teams on L2/L3 escalations—communicating findings, driving root-cause analysis, and resolving complex packaging, deployment, upgrade, and runtime issues across heterogeneous Kubernetes environments. • Drive issues to resolution by reproducing problems locally, isolating root causes, and coordinating fixes with engineering—then documenting learnings in crisp RCAs that become actionable improvements • Build and maintain diagnostics tooling including support bundles, health checks, environment validators, and "what changed?" helpers that make future troubleshooting 10x faster • Own the test automation infrastructure roadmap, improving CI stability, reducing flaky tests, and creating reproducible integration/e2e environments that catch issues before customers do • Establish and maintain performance baselines and regression tests that serve as actionable gates, helping teams catch scale and latency issues early • Improve installation and upgrade robustness by identifying recurring failure modes and eliminating them through product changes, automation, and guardrails • Write production-quality code in Python, Go, or Rust for internal tooling and product improvements that directly enhance reliability • Close the reliability feedback loop by systematically turning field issues into better tests, observability, documentation, and product defaults—measuring success through reduced time-to-resolution and fewer repeat incidents
Building secure infrastructure for seamless cross-chain digital asset movement.
• Build and maintain CI/CD pipelines for multi-service environments • Manage cloud infrastructure (AWS / GCP) and deployment workflows • Design and implement monitoring, logging, and alerting systems • Ensure system reliability, scalability, and high availability • Manage containerized environments using Docker • Optimize infrastructure cost and performance • Support blockchain node infrastructure and RPC reliability
Thumbtack helps millions of people confidently care for their homes. Thumbtack is the one app you need to take care of and improve your home — from personalized guidance to AI tools and a best-in-class hiring experience. Every day in every county of the U.S., people turn to Thumbtack to complete urgent repairs, seasonal maintenance, and bigger improvements. We help homeowners know which projects to do, when to do them, and who to hire from our growing community of 300,000 local service businesses. If making an impact inspires you, join us. Imagine what we’ll build together.
• Design, create, and maintain software and systems to improve the availability, scalability, and efficiency of Thumbtack's services • Set the architectural direction of infrastructure and platform services while supporting the engineering organization • Design and implement tools and processes used for deployment, change, service, and infrastructure management • Troubleshoot and debug critical systems throughout the SDLC • Contribute to the evolution and performance of capabilities we provide to engineering as a platform organization • Capacity planning and demand forecasting, anticipating performance bottlenecks • Participate in rotating on-call duties
What makes CellPoint Digital a leader in the payment landscape isn’t just our technology - it’s our people and how we work together. We’ve built a global community where diverse talents and perspectives unite to create innovative solutions. When you join us, you become part of something bigger: a collaborative culture that crosses borders and disciplines, bringing out the best in every team member to deliver breakthrough results for our clients and partners. Together, we are transforming the payments industry - challenging, supporting, and inspiring one another in the process.
Role Description Join us as a Senior Site Reliability Engineer on our mission to turn payments into possibilities! The Site Reliability Engineering (SRE) team ensures the reliability, availability, scalability, and performance of a mission-critical payment orchestration platform. The platform operates in a high-volume, API-driven environment and supports integrations with multiple payment service providers. - SREs collaborate closely with the engineering, product, security, and operations teams to maintain resilient systems. - Lead incident response, implement release management processes, and continuously improve platform stability through automation, observability, and infrastructure best practices. - The Senior SRE provides technical leadership across the SRE function, owns reliability strategy, release management, major incident leadership, and security posture initiatives. - All SRE roles include participation in a 24×7 on-call rotation. Please note: We are currently seeking candidates who are available to join immediately or within a short notice period . Applications from candidates with extended notice periods will not be considered at this time. Qualifications - Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field. - 6+ years of experience in SRE, DevOps, or platform reliability roles. - Deep expertise in GCP, including Kubernetes (GKE), CloudSQL, Spanner, networking, and IAM. - Experience supporting high-volume, mission-critical payment or financial systems. - Advanced hands-on experience with Terraform and infrastructure automation. - Proven ability to lead major incidents and influence cross-functional teams. - Excellent communication, documentation, and stakeholder management skills. Requirements - You are eager to bring your unique talents and authenticity to the CellPoint Digital community. - You're constantly curious and a lifetime learner. - You have excellent communication and relationship-building skills. - You enjoy leading and supporting cross-functional initiatives and projects in a team where you are empowered and accountable. - You thrive in a fast-paced environment and the challenge of managing multiple projects simultaneously while prioritizing high-return work. - You approach challenges with a solution-oriented mindset. - You are able to thrive in a ‘remote first’ arrangement with a distributed organization in multiple time zones. Benefits - Opportunity to be an innovator, challenge the status quo, and redefine the payments category. - Competitive salary in a fast-growing start-up. - Medical insurance with coverage for dependents (parents, spouse, children). - Rewards & Recognition system. - Opportunity for personal and professional growth in a dynamic industry.
2,133more opportunities are still waiting for you.Log in now and take your next shot before someone else does.
AWS, Cloud, Python, Terraform, Linux, EC2