H&R Block logo
H&R Block

With expert guidance, upfront pricing, and more ways to file, it’s #BetterWithBlock.

Lead Database Reliability Engineer – Small Business Technologies

Location

Missouri

Posted

4 days ago

Salary

$117.7K - $188.3K / year

Seniority

Senior

Job Description

Lead Database Reliability Engineer – Small Business Technologies

H&R Block

• Own end-to-end reliability of critical database systems (availability, performance, durability) • Define and maintain SLIs/SLOs for database services • Lead incident response and root cause analysis for database-related outages • Design and implement failover, replication, and disaster recovery strategies • Automate database provisioning, scaling, backups, and recovery processes • Implement Infrastructure-as-Code (IaC) and GitOps approaches for database environments • Drive reduction of manual operational work through tooling and self-service capabilities • Establish reliability engineering practices (error budgets, postmortems, continuous improvement) • Build and maintain database observability (metrics, logs, traces) • Proactively monitor system health and performance bottlenecks • Optimize query performance, indexing strategies, and resource utilization • Establish capacity planning and forecasting models • Partner with platform engineering to standardize database patterns and tooling • Support cloud-native database architectures (managed services, distributed systems) • Evaluate and recommend database technologies and storage solutions • Ensure systems are designed for resilience, scalability, and cost efficiency • Own backup strategies, validation, and restore testing • Lead disaster recovery planning and execution readiness • Ensure data integrity, consistency, and compliance with governance standards • Work with engineering teams to improve data access patterns and reliability • Partner with security and compliance teams on data protection requirements • Influence application design to align with database reliability best practices • Act as the technical lead for database reliability across teams • Define standards, best practices, and operating models for DBRE • Lead cross-team initiatives to improve reliability and reduce systemic risk • Mentor senior engineers and contribute to talent development

Job Requirements

  • 8–12+ years in database engineering, SRE, or infrastructure engineering
  • Hands-on experience with production-grade relational and/or NoSQL databases (e.g., PostgreSQL, MySQL, SQL Server, DynamoDB, etc.)
  • Experience operating systems at scale in cloud environments (AWS, Azure, or GCP)
  • Proven ownership of high-availability, mission-critical systems
  • Strong knowledge of: Database internals, replication, and recovery mechanisms, Performance tuning and optimization techniques, Distributed systems and fault-tolerant architectures
  • Experience with: Observability tools (Datadog, Prometheus, Grafana, etc.), Automation tools (Terraform, Ansible, CI/CD pipelines), Containerization and orchestration (Docker, Kubernetes)
  • Familiarity with data lifecycle management and storage optimization
  • Demonstrated ability to lead through influence in cross-functional environments
  • Strong incident management and problem-solving skills under pressure
  • Ability to communicate complex technical concepts to leadership stakeholders
  • Experience defining engineering standards and operational practices

Benefits

  • Qualifying associates can enroll themselves and/or their eligible dependents in medical and prescription drug coverage
  • Participate in the H&R Block Retirement Savings Plan (401(k) Plan)
  • Employee Assistance Program
  • (virtual) fitness center programs
  • Associate discount program
  • Automatically enrolled in Business Travel Accident Insurance
  • Receive Associate Tax Prep benefit

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Guild logo

Manager, DevOps

Guild

At Guild, we unlock opportunity for America’s workforce through education, skilling, and career mobility.

DevOps Engineer4 days ago
Full TimeRemoteTeam 1,001-5,000Since 2015H1B Sponsor

Role Description Guild is seeking an Engineering Manager to lead our DevOps team. This role involves overseeing our cloud infrastructure, automation frameworks, CI/CD pipelines, and platform services. The ideal candidate will possess expertise in Infrastructure as Code (IaC), security, automation, networking, and cloud governance. The Engineering Manager will collaborate with engineering, security, and product leaders to ensure DevOps initiatives align with business goals and advance Guild’s cloud platform. Role Responsibilities - People Leadership & Team Growth - Develop, mentor, and grow a high-performing DevOps team, fostering technical culture and career progression. - Build redundancy and skill diversity within the team to ensure knowledge sharing. - Champion and set performance expectations, guiding engineers in accountability, ownership, and execution. - Lead through ambiguity and change, maintaining focus on long-term goals. - AWS Cloud Infrastructure & Platform Ownership - Oversee and ensure Guild’s cloud infrastructure is scalable, cost-efficient, and secure. - Own and optimize Guild’s multi-account AWS cloud strategy for secure, scalable, and efficient networking and service connectivity. - Drive IaC best practices to align cloud resources with security and compliance. - Ensure system reliability, security, and performance, while proactively addressing operational risks and technical debt. - Automation, CI/CD & Observability Enablement - Ensure fast, secure, and scalable build and deployment pipelines for all engineering teams. - Improve DevOps automation tooling to enable efficient service deployment, monitoring, and management. - Enhance system observability across multi-account AWS environments. - Optimize workflows and deployment strategies to reduce manual intervention and ensure seamless software delivery. - Strategic Alignment & Cross-Team Collaboration - Partner with Product, Security, and Engineering Leadership to align DevOps initiatives with business goals. - Ensure DevOps services adoption and drive best practices for infrastructure automation. - Align infrastructure investments with business priorities. - Support compliance, security, and cloud governance efforts. - Operational Excellence & Continuous Improvement - Lead postmortems, incident response, and a culture of continuous improvement. - Own and evolve team processes, ensuring clear priorities and accountability. - Monitor cloud costs and drive FinOps best practices to ensure cost-efficiency. Qualifications - 5+ years of experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles, with at least 3+ years in a leadership or management position. - Deep expertise in AWS cloud services, including IAM, VPC networking, EC2, Lambda, RDS, S3, and security best practices. - Hands-on experience with Infrastructure as Code (IaC) using AWS CDK, CloudFormation, and automation tooling. - Strong background in CI/CD pipeline design and implementation, including experience with GitHub Actions or similar deployment automation tools. - Experience managing multi-account AWS cloud environments, including networking, security controls, IAM governance, and cost optimization. - Proficiency in observability and monitoring solutions (e.g., Datadog, CloudWatch, Prometheus, OpenTelemetry) to drive system reliability and operational excellence. - Strong scripting and automation skills in Python, Bash, or another high-level language to develop internal DevOps tooling. - Proven track record of leading and mentoring engineering teams, setting clear goals, and fostering a high-performance, inclusive, and collaborative DevOps culture. - Experience working in an Agile, iterative development environment, ensuring infrastructure and automation solutions are aligned with engineering needs and business priorities. - Demonstrated ability to lead cross-functional collaboration, working with engineering, security, compliance, and product teams to drive platform-wide improvements. Requirements - Competitive total compensation package, including a base salary of $190,000 - $220,000, and stock options. - Compensation offered will be based on a combination of factors such as experience, competencies, and internal equity. Benefits - Access to low-cost, high-quality health care options through Collective Health and Kaiser (due to coverage limitations, Kaiser is currently only available in CA & CO). - Access to a 401k to help save for the future. - Vacation policy to rest and recharge. - 8 days of fully-paid sick leave, to take the time to heal and or recover. - Family-friendly benefits, including 12 weeks of parental leave for non-birthing parents and 18-20 weeks for birthing parents; 2-week ramp-up period for when employees return from a leave of 6 weeks or more; as well as employer-paid short-term and long-term disability, employer-sponsored life insurance, fertility and caregiving benefits. - Well-rounded wellness benefits including free and low-cost mental health resources and financial wellbeing support services. - Education benefits and tuition assistance to help your future development and growth. - Our home and headquarters in Denver, but some of our roles allow for remote work, allowing us to reach the best talent across the US.

United States
$190K - $220K / year
Join Creative Tech logo

Senior DevOps Analyst

Join Creative Tech

Criamos projetos de software com propósito e olhar criativo.

DevOps Engineer4 days ago
Full TimeRemoteTeam 51-200Since 2010H1B No Sponsor

• Administer and maintain LAN, WAN, Wi‑Fi, VPN, firewalls, communication links, and corporate connectivity. • Configure, analyze, and troubleshoot switches, routers, wireless controllers, firewalls, and network security solutions. • Support the implementation of security policies, network segmentation, firewall rules, access control, and hardening of network assets. • Monitor availability, performance, capacity, and events related to network and security infrastructure. • Perform analysis and resolution of incidents, problems, and changes related to connectivity and infrastructure security. • Support vulnerability management processes, rule reviews, risk identification, and improvements to the security posture. • Produce technical documentation, network diagrams, operational procedures, and improvement recommendations. • Collaborate with infrastructure, security, and operations teams, vendors, and business units to ensure service continuity and quality.

Brazil
Full TimeRemoteTeam 11-50H1B No Sponsor

• Drive the reliability of Epic's infrastructure—set and track SLOs/SLIs, reduce toil, and engineer out recurring instability. • Build and operate the cloud infrastructure and container platform for high availability, scalability, and cost efficiency—including workload scheduling, autoscaling, networking, and graceful failure handling. • Maintain and improve CI/CD pipelines for fast, safe delivery across engineering teams. • Own and evolve the observability stack—metrics, logs, traces, dashboards, and alerts. • Manage infrastructure as code across the organization, with a focus on consistency, change safety, and reproducibility. • Own platform security practices—including secrets management, IAM policies, and network segmentation. • Support compliance-aware infrastructure practices—including vulnerability management, access reviews, audit-evidence flows, and incident-response readiness. • Participate in a frequent on-call rotation; drive incident response, blameless post-mortems, and follow-through on systemic fixes. • Partner with product and data engineering teams to troubleshoot platform issues and guide developers on infrastructure best practices.

United States
$160K - $200K / year
Alpaca logo

Senior DevOps Engineer

Alpaca

Developer APIs for stocks and crypto trading, investing apps, and embedded fintech.

DevOps Engineer4 days ago
Full TimeRemoteTeam 201-500H1B No Sponsor

• Design and evolve our cloud architecture on GCP - networking, interconnects, IAM and high-availability topology - and express it entirely as code with Terraform, following GitOps as a first principle. • Build and own the CI/CD pipelines that plan, review, test and safely apply IaC changes - Policy-as-Code guardrails, drift detection and progressive rollout so infrastructure changes ship as confidently as application code. • Advance Platform-as-a-Product: build self-serve capabilities and paved paths so engineers can provision what they need, through a golden path rather than a hand-off. • Strengthen our observability stack - metrics, logs, traces and alerting across Prometheus, Thanos, Grafana, Loki, Tempo and Alertmanager - so the platform is easy to run and reason about. • Operate our GKE clusters and the infrastructure services that run on them - Helm-packaged workloads, message brokers (RabbitMQ, IBM MQ) and data stores. • Participate in our Follow-The-Sun on-call model: watch and triage alerts, join and declare incidents, lead structured debugging and escalation, and drive blameless post-mortems and the post-actions that actually close the loop. • Embed SRE practices - SLIs/SLOs and error budgets, capacity planning - into how Core Infrastructure builds and operates, working closely with our SRE function.

Japan