Job Closed

This listing is no longer active.

Rocket Mortgage logo
Rocket Mortgage

Rocket Mortgage® is the home loan experience designed for you. NMLS #3030

Director, Data Reliability Engineering

DevOps EngineerDevOps EngineerFull TimeRemoteLeadTeam 10,001+Since 1985H1B SponsorCompany SiteLinkedIn

Location

Michigan

Posted

63 days ago

Salary

$128.5K - $276K / year

Seniority

Lead

Bachelor Degree10 yrs expEnglishAWSCloud

Job Description

Director, Data Reliability Engineering

Rocket Mortgage

• Lead Engineering teams responsible for improving the reliability, observability, recoverability, and operational maturity of enterprise data platforms • Define reliability standards for databases, data warehouses, pipelines, jobs, storage, access patterns, and supporting infrastructure • Establish operating expectations for monitoring, alerting, logging, incident response, change management, backup/recovery, disaster recovery, patching, access controls, service ownership, and operational readiness • Create metrics that measure platform health, data freshness, data quality, recovery readiness, incident trends, operational risk, compliance alignment, and business impact • Lead current-state assessments of systems, data flows, operational processes, observability, access patterns, and reliability gaps • Convert assessment findings into executable roadmaps that improve platform stability, data trust, security alignment, and operational predictability • Support migration and modernization programs involving on-premise platforms, AWS, Snowflake, and related enterprise data systems • Build durable operating mechanisms, including reliability reviews, service health reviews, incident reviews, operational readiness reviews, risk reviews, roadmap reviews, and executive reporting • Develop senior technical talent and create the leadership structure required to scale Data Reliability Engineering over time

Job Requirements

  • 10+ years of experience in data infrastructure, database engineering, data platform engineering, cloud infrastructure, site reliability engineering, or related technical disciplines
  • 5+ years of experience leading engineering teams responsible for production systems, databases, data platforms, infrastructure platforms, or reliability engineering
  • Strong understanding of enterprise data infrastructure, including databases, data warehouses, pipelines, storage, compute, backup/recovery, resiliency, and production operations
  • Experience improving reliability practices across complex production environments, including observability, monitoring, incident response, change management, disaster recovery, and lifecycle management
  • Experience establishing service health metrics, data reliability metrics, operational maturity indicators, and executive-level reporting
  • Strong understanding of enterprise security, compliance, access management, auditability, operational controls, and infrastructure standards
  • Proven ability to create structure in ambiguous environments, set clear priorities, influence across teams, and translate technical reliability work into business outcomes

Benefits

  • Perks and health benefits for you and your family
  • Support for individual needs
  • Peace of mind with our offerings

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Hunt St logo

Senior DevOps Engineer

Hunt St

We help Aussie companies find top 3% remote talent in the Philippines & Nepal for a single finder's fee.

DevOps Engineer63 days ago
Full TimeRemoteTeam 1-10H1B No Sponsor

Role Description We are seeking an experienced and highly skilled Senior DevOps Engineer to join our engineering team. This role is critical to the development and deployment of our infrastructure, ensuring robust CI/CD pipelines, infrastructure as code (IaC), cloud environment optimization, and seamless collaboration across development, QA, and operations teams. The ideal candidate is passionate about automation, performance, scalability, and reliability. Key Responsibilities - Maintaining and improving the resiliency of our core applications and our hybrid infrastructure platform - Providing continued improvement to the platform infrastructure through automation and standardisation - Providing complementary skills and expertise to the teams and continuously learning from peers and seniors - Ensuring that all of our core services are up to date and security patched - Working closely with development teams to ensure applications are configured for security, efficiency and scalability Qualifications - Bachelor's degree in Computer Science, Information Technology, or a related field - 5+ years of experience in DevOps, Systems Engineering, or a related field - Linux native; if you do not use Linux as your preferred OS this may not be the role for you - Great communication skills (verbal and written) - Strong experience with the following: - Linux administration - Bash scripting - Kubernetes - Docker - AWS - Good knowledge of networking, DNS, load balancing and CDN's Preferred Qualifications - Experience with Terraform and Ansible - Experience working on and supporting container-based CI/CD pipelines - Keen interest in SecOps practices - AWS Certifications (we will fully support any AWS certification you are seeking) - Experience configuring observability platforms for monitoring and alerting (including Prometheus and New Relic) - Experience with Hashicorp vault, Redis, RabbitMQ or MSSQL - Experience with any programming languages (i.e. Node JS, PHP, Typescript, Python) Work Arrangement & Expectations This is a remote role that will be set up as an independent contractor engagement. To ensure alignment and transparency, successful candidates will be expected to: - Be available for meetings and collaboration during core [AEST or PHT] business hours - Disclose any existing ongoing roles or client work - Reflect this engagement on their LinkedIn profile (clearly marked as “Independent Contractor”)

Philippines
A$4K / month
Full TimeRemoteTeam 1-10Since 2024H1B No Sponsor

• Design, deploy, and manage containerized workloads using Amazon ECS (Elastic Container Service) and Amazon EKS (Elastic Kubernetes Service). • Build and maintain CI/CD pipelines to automate software delivery workflows. • Develop and manage Docker container images, registries (ECR), and container lifecycle best practices. • Implement Infrastructure as Code (IaC) using tools such as Terraform, CloudFormation, or CDK. • Monitor, troubleshoot, and optimize cloud infrastructure performance, availability, and cost. • Enforce security best practices across containerized environments (IAM roles, network policies, secrets management). • Collaborate with software engineers to containerize applications and migrate workloads to ECS/EKS. • Manage Kubernetes cluster configurations, namespaces, Helm charts, and service mesh integrations. • Define and maintain observability standards using tools like CloudWatch, Prometheus, Grafana, or Datadog. • Participate in on-call rotations and incident response processes.

United Kingdom
Job Closed
Akka (formerly Lightbend) logo

Lead Site Reliability Engineer

Akka (formerly Lightbend)

Responsive by Design, Akka apps are elastic, agile, and resilient.

DevOps Engineer63 days ago
Full TimeRemoteTeam 51-200Since 2011H1B No Sponsor

• Own Service Level Objectives/Service Level Indicators (SLOs/SLIs) and error budgets across multi-cloud clusters (EKS, GKE, AKS); drive blameless post-mortems and systemic remediation. • Lead capacity planning with our customers, cluster lifecycle management, and Kubernetes and database upgrade cycles. • Define and enforce runbooks, on-call rotations, and escalation paths for the wider engineering organisation. • Own and evolve the IaC layer: Helm charts, Crossplane compositions, and FluxCD GitOps pipelines. • Design and maintain cloud-resource provisioning workflows that span all three cloud providers, with consistent policy controls. • Architect and operate connectivity patterns: AWS PrivateLink / Transit Gateway, GCP NCC, Azure VNet Peering, and cross-region ingress with Contour/Envoy. • Maintain and evolve the Linkerd service mesh for mTLS, workload identity (OIDC), and zero-trust authorisation policies. • Drive PKI hygiene with cert-manager: root/intermediate CA rotation, ACME certificate lifecycle, and secret management via KMS-backed Kubernetes vaults. • Own the observability stack: Prometheus, Cortex (multi-tenant metrics), OpenTelemetry sidecars, centralised log pipelines, and Groundcover / Grafana dashboards. • Establish alerting standards and SLO-based alerting rules; ensure distributed traces are actionable across JVM, Rust, and Go workloads. • Actively participate in on-call and lead the technical response for platform-level incidents. • Set engineering standards and review infrastructure changes across the team. • Partner with Security, Product, and Application Engineering to translate reliability requirements into platform capabilities. • Grow a team of 3–5 SREs through code review, architecture sessions, and career conversations.

United States
Job Closed
Full TimeRemoteTeam 501-1,000Since 2005H1B No Sponsor

• Lead Reliability Engineering for User Experience • Drive reliability, scalability, and operational excellence for critical user facing systems and services. Improve performance and resiliency across APIs, content delivery, feed generation, search, messaging, and real-time experiences. • Partner with product and infrastructure engineering teams to design systems that remain highly available and performant under massive global load. Guide architectural decisions around failover, redundancy, graceful degradation, traffic management, and capacity planning. • Identify systemic risks and reliability bottlenecks across services, dependencies, deployments, and infrastructure. Build proactive mitigation strategies and drive engineering improvements that reduce incidents and improve service health. • Eliminate repetitive operational work through automation and tooling. Build systems that improve deployment safety, incident response, remediation workflows, and reliability guardrails • Lead complex incident response efforts across engineering teams. Drive blameless postmortems, identify root causes, and ensure sustainable long-term fixes are implemented. • Define and champion best practices around reliability engineering, SLIs/SLOs, capacity management, release engineering, and operational maturity across the company. • Provide technical leadership and mentorship to engineers across SRE and software engineering teams. Help shape reliability culture and raise the operational excellence bar across the organization.

United Kingdom