Senior DevOps Engineer, Applications

Location

California + 4 moreAll locations: California | Colorado | New York | Oregon | Washington

Posted

58 days ago

Salary

$131.3K - $175K / year

Seniority

Senior

Job Description

Senior DevOps Engineer, Applications

Endeavor

• Own and evolve Azure infrastructure using Terraform — networking, compute, managed services, and secrets • Manage Kubernetes clusters and Helm-based deployments for a distributed microservices platform • Build and maintain CI/CD pipelines (GitHub Actions / Azure DevOps) that keep deployments fast and safe • Define and improve observability: metrics, logging, alerting, and tracing across services • Lead incident response — drive resolution, write postmortems, and follow through on reliability improvements • Manage PostgreSQL infrastructure: backups, replication, failover, and query performance • Collaborate closely with engineers to improve local development environments and deployment workflows • Enforce security best practices: secrets management, network policies, access controls, and vulnerability scanning

Job Requirements

  • Solid DevOps or platform engineering experience, typically 5+ years
  • Strong hands-on experience with Azure (AKS, Azure AD, managed databases, blob storage, networking)
  • Terraform experience — writing, organizing, and maintaining infrastructure as code at scale
  • Kubernetes and Helm experience managing production workloads
  • Experience building and operating CI/CD pipelines (GitHub Actions, Azure DevOps, or similar)
  • PostgreSQL experience — operational management, migrations, and performance tuning
  • Solid understanding of observability tooling (Datadog, Grafana, OpenTelemetry, or equivalent)
  • Experience with incident response and driving reliability improvements from postmortems
  • Docker and container image management experience
  • Familiarity with secrets management tools (Azure Key Vault, Vault, or equivalent)
  • Comfortable working in environments where systems and processes are still evolving.

Benefits

  • health care
  • retirement
  • vacation and other paid time off
  • growth and developmental opportunities

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Full TimeRemoteTeam 10,001+Since 1993H1B Sponsor

• Working on building tools to improve the SRE Observability. • Be part of the Kubernetes migration journey with VMI setup and problem solving. • Rapidly debug and triage incidents and user-reported issues • Taking ownership of automating, scripting, and tooling of new/existing scripts to help the team achieve 100% automation of daily tasks • Support services before they go live through activities such as system design consulting, developing software platforms and frameworks, capacity management and launch reviews. • Be part of an on call rotation to support production systems

California
$168K - $270.3K / year
Job Closed
Oowlish logo

Azure DevOps Engineer – Cloud & AI Delivery

Oowlish

We make innovation simple, convenient and right...we just make it HAPPEN

DevOps Engineer58 days ago
Full TimeRemoteTeam 51-200Since 2017H1B No Sponsor

• Design, build, and maintain Azure-based cloud infrastructure • Develop and optimize CI/CD pipelines for multiple engineering squads • Support cloud-native application deployments and automation • Improve software delivery processes and engineering efficiency • Collaborate with developers, architects, and technical leadership • Monitor, troubleshoot, and optimize cloud environments • Implement infrastructure-as-code and automation best practices • Support security, reliability, and scalability initiatives

Brazil
eTelligent Group LLC logo

Platform Operations and Site Reliability Lead

eTelligent Group LLC

Over the past 15 years, eTel has delivered essential solutions for the federal government by securing and managing data, providing scalable identity access, modernizing legacy systems, and building high-performance platforms. By integrating new technologies and ensuring reliable operations we help agencies stay prepared for future challenges. eTel offers integrated CMMI Level 3 processes, tools, and techniques with innovative, cost-efficient, and secure solutions to address complex challenges. eTel holds ISO 9001:2015, ISO/IEC 27001:2013, and ISO/IEC 20000-1:2018 certifications. Offers dedicated subject matter experts (SMEs) and thought leaders that possess a deep understanding of customers’ environments and challenges.

DevOps Engineer58 days ago
Full TimeRemoteTeam 51-200

Role Description The Platform Operations and Site Reliability Lead is responsible for ensuring the reliability, availability, performance, scalability, and operational excellence of the Enterprise Data Platform. The Operations Lead oversees 24x7 platform operations, observability, incident response, disaster recovery, performance optimization, and AI enabled operational automation across AWS and Databricks environments. Key Responsibilities - Lead operations and maintenance activities supporting AWS cloud infrastructure and Databricks E2 services. - Manage observability frameworks including monitoring, logging, tracing, and alerting. - Implement Site Reliability Engineering practices including SLIs, SLOs, error budgets, and reliability metrics. - Coordinate incident response, root cause analysis, and service restoration activities. - Develop operational runbooks, playbooks, and automated remediation procedures. - Lead disaster recovery planning, testing, backup validation, and continuity activities. - Support AI driven operational intelligence and predictive monitoring capabilities. - Track and report service levels, uptime metrics, and operational performance indicators. Qualifications - Minimum 8 years managing enterprise production environments. - Minimum 5 years supporting AWS cloud operations. - Experience supporting Databricks, analytics platforms, or enterprise data environments. - Experience implementing enterprise monitoring, observability, and Site Reliability Engineering practices. Preferred Certifications - AWS Certified SysOps Administrator - AWS Solutions Architect Associate - Databricks Platform Administrator Citizenship - US Citizen (MUST) Security Clearance - Must be eligible to possess MBI (IRS Background Investigation) clearance. - Active IRS MBI clearance is preferred. Commitment to Diversity eTelligent Group provides equal employment opportunities (EEO) to all applicants without regard to race, color, religion, gender, sexual orientation, gender identity, nations origin, age, disability, genetic information, marital status, amnesty, status as a covered veteran, and any other characteristic provided in accordance with applicable, federal, state and local laws.

United States
Job Closed
Full TimeRemoteTeam 51-200

Role Description Ocient is searching for an experienced Site Reliability Engineer with strong problem-solving skills and a passion for solving hard problems to help maintain and expand Ocient's "as a service" offering of its cutting-edge data warehouse. - Support the design and operations of Ocient's hosted database and related services — including message queues and storage systems — ensuring high availability, performance, and efficiency. - Design and maintain monitoring, log centralization, and alerting for all services to facilitate observability and incident management. - Automate deployment and configuration of Linux-based servers, including the OS and the numerous applications that compose our hosted offerings. - Develop and maintain rigorous security practices to protect our applications and customer data. - Assist with automation of testing pipelines for the Ocient DB and monitoring of test infrastructure. Qualifications - 3+ years of experience in system administration in production environments. - Scripting experience with Bash, Python, or other languages. - Experience with system and software monitoring and alerting tools, such as the ELK stack, Graylog, InfluxDB, Prometheus, Zabbix, Grafana, Dynatrace, or others. - Experience with configuration management software such as Ansible, Puppet, or Chef. - Experience with data archiving, backup and disaster recovery. - Continuous Integration / Continuous Deployment experience with Jenkins, Gitlab CI or others. - Experience with source control tools like Git. - Ability to work flexible hours and serve in on-call rotations. Requirements - Knowledge of OWASP principles for application security. - Experience with server/system virtualization and containerization technologies e.g., ProxMox, KVM, VMware. - Experience with SQL and Database Administration. - Experience managing and operating cloud infrastructure (e.g. AWS, GCP, Azure). - Experience with SSAE18 SOC2 Compliance. - Experience with networking administration, including VPN, proxy, DNS, and firewall configuration.

United Kingdom
£74K - £90K / year