Data Reliability Engineer
Location
United States
Posted
10 days ago
Salary
0
Seniority
Senior
Job Description
Data Reliability Engineer
Vytalize Health
• Own and continuously improve the reliability of data pipelines across ingestion, transformation, and delivery layers, ensuring data is accurate, complete, and delivered on schedule. • Establish and maintain data reliability standards, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs) for both upstream ingestion and downstream data delivery. • Design, implement, and maintain comprehensive monitoring, logging, and observability frameworks for data pipelines, datasets, and data services with clear visibility into freshness, volume, schema changes, and data quality. • Design and implement data quality testing and validation frameworks — establishing test cases, golden datasets, and regression tests to detect quality issues early. • Lead incident response for data reliability issues, including detection, triage, communication, root cause analysis, and post-incident remediation with documented corrective actions. • Drive improvements in pipeline resiliency through retry strategies, backfills, idempotency, schema enforcement, and safe deployment practices.
Job Requirements
- Bachelor's degree in Computer Science, Engineering, Information Systems, or related field, or equivalent professional experience.
- 5+ years of experience working with data platforms, data pipelines, or distributed data systems in production environments.
- Demonstrated experience improving reliability, observability, or operational quality of data systems with measurable SLI/SLO/SLA improvements.
- Hands-on experience supporting both data ingestion pipelines and downstream data consumption or delivery patterns.
- 1+ years of hands-on experience with machine learning-based monitoring, anomaly detection, or AI-assisted observability tools.
- Demonstrated experience with data quality testing, validation frameworks, and quality metrics definition.
- Proficiency in Python and SQL, with experience building or supporting production-grade data pipelines.
- Strong understanding of modern data architectures, including data lakehouse patterns and multi-layer (bronze/silver/gold) data models.
- Experience with cloud-based data platforms (AWS, Databricks, or similar).
Benefits
- Health insurance
- 401(k) matching
- Flexible work arrangements
- Professional development opportunities
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Infrastructure / Systems Operations Engineer
ZignalyPassionate individuals at Zignaly are building a crypto investment platform to level the playing field for everyone.
• Operate and maintain our Linux, AWS, and blockchain-node infrastructure on a day-to-day basis. • Own routine administration, monitoring, and troubleshooting of services and networking. • Contribute to our infrastructure-as-code work with Terraform, growing your ownership over time. • Support blockchain node operations across our ecosystem. • Automate recurring tasks and diagnostics through scripting. • Help keep our systems secure, patched, and running smoothly.
Role Description As a Senior DevOps Engineer, you will work closely with Product, Engineering, and AI teams to shape our infrastructure strategy, design resilient cloud architectures, and ensure our platforms are secure, scalable, and high-performing. You will play a key role in bringing AI systems into production, enabling reliable delivery, strong observability, and operational excellence across our products and internal systems. - Design and operate secure, scalable, and high-quality infrastructure that supports modern applications and advanced AI workloads. - Build and maintain robust automation across CI/CD pipelines, infrastructure provisioning, and operational processes to improve reliability and minimize manual effort. - Integrate AI-driven solutions into operational workflows to enhance efficiency, detect anomalies, and accelerate delivery. - Apply strong systems engineering practices, including monitoring, incident management, performance optimization, and capacity planning. - Establish and uphold DevOps best practices, ensuring reproducibility, testing, documentation, and operational excellence. - Communicate technical decisions clearly and collaborate cross-functionally to support predictable delivery and effective problem-solving. - Provide mentorship and technical leadership, raising the level of platform engineering, DevOps maturity, and overall engineering quality across the organization. Qualifications - 6+ years of progressive experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Engineering. - Strong, hands-on experience across multi-cloud environments (AWS, GCP, Azure), including expertise in networking, compute, storage, security, and cost optimization. - Deep expertise in containerization and orchestration and extensive experience with Infrastructure as Code (IaC) (e.g., Terraform, Pulumi, CloudFormation). - Experience supporting or deploying AI/ML workloads (e.g., model inference, vector databases, GPU workloads), or strong familiarity with the infrastructure requirements for these systems. - Proven ability to design, build, and operate highly reliable, scalable production systems utilizing advanced Zero-Downtime Deployment Patterns (e.g., Blue/Green, Canary, progressive delivery, Preview Environments). - Expertise in modernizing deployments via GitOps practices (e.g., ArgoCD, Flux) and building Self-Service Developer Platforms that enable engineering efficiency (e.g., environment automation, internal tooling). - Experience implementing and managing Multi-Cloud API Gateways and Edge Routing solutions. - Strong background in platform security, including secrets management, Identity and Access Control (IAM), and Runtime/Security Hardening. - Solid understanding and practical experience with modern observability stacks. - Excellent communication and collaboration skills with a proven ability to describe complex infrastructure decisions clearly and a background in mentoring engineers and driving improvements in engineering practices. - Familiarity with modern programming languages like Node.js, NestJS, and Python is highly desirable for extending DevOps capabilities or integrating tooling.
Role Description Gestalte moderne Infrastruktur mit uns. Werde Teil der Contensi Software GmbH. Zur Verstärkung unseres Teams suchen wir eine:n Senior DevOps Engineer, der/die mit fundierter Erfahrung, Weitblick und Leidenschaft für Technologie den Unterschied macht. - Du arbeitest eng mit unseren Kunden aus Industrie, öffentlichem Sektor und Wissenschaft an anspruchsvollen Infrastrukturprojekten, von der Planung bis zum stabilen Betrieb. - Du übernimmst die Architektur, Einrichtung und Weiterentwicklung von Cloud- und OnPrem-Infrastrukturen, darunter Public-Cloud-Umgebungen wie AWS, Azure, GCP sowie Plattformen wie Kubernetes und OpenShift. - Du betreibst und wartest bestehende Systeme und Rechenzentrumsinfrastrukturen, inklusive Hybrid- und Multi-Cloud-Setups. - Du analysierst bestehende Architekturen und unterstützt bei der Migration und Modernisierung von Legacy-Umgebungen. - Du richtest GitOps- und CI/CD-Workflows ein (z. B. mit ArgoCD, GitLab CI, Azure DevOps) und begleitest unsere Kunden bei der Automatisierung ihrer Entwicklungs- und Betriebsprozesse. - Du berätst zu aktuellen Technologien und Best Practices, insbesondere im Bereich Containerisierung, Cloud-native Tools und Infrastructure-as-Code (Terraform, Ansible). - Du arbeitest konzeptionell mit an Themen wie Netzwerkdesign, Storage-Architekturen (z. B. Ceph, ODF, NFS) und Security-Standards (BSI, ISO27001, IAM, PKI). - Bei Wunsch betreust du Kunden auch vor Ort in Deutschland – und unterstützt unsere Kolleg:innen durch dein Fachwissen im Team. Qualifications - Mindestens 6 Jahre Berufserfahrung im Bereich DevOps, Infrastruktur oder Site Reliability Engineering. - Erfahrung mit Cloud-Plattformen: AWS, Azure, GCP. - Sehr gute Kenntnisse in Kubernetes und OpenShift. - Praktische Erfahrung mit einem oder mehreren CI/CD-Tools: Jenkins, GitLab CI, GitHub Actions, Azure DevOps o. ä. - Kenntnisse in Logging- und Monitoring-Systemen wie ELK, Datadog, Prometheus, New Relic o. ä. - Sehr gute Kenntnisse in Infrastructure as Code (Terraform, Ansible, Puppet). - Solides Netzwerkverständnis (TCP/IP, DNS, Routing, Firewalls etc.). - Erfahrung mit Storage-Systemen wie Ceph, MinIO etc. - Know-how in Security-Standards & Compliance (z. B. BSI, ISO27001, PKI, IAM). - Fundierte Kenntnisse in mindestens einer Programmiersprache: z. B. Java, Python, Go, Node.js oder TypeScript. - Du bist motiviert, dich kontinuierlich weiterzuentwickeln und neue Technologien zu erlernen. - Fliessende Deutsch- und Englischkenntnisse. Benefits - Spannende Kundenprojekte mit technologischer Tiefe. - Remote-Arbeit mit gelegentlichen vor Ort Besuchen beim Kunden. - Gestaltungsfreiraum - bring deine Ideen ein, übernimm Verantwortung und wachse mit uns. - Weiterbildungsbudget für Zertifikate, Konferenzen und Fachliteratur. - Modernstes Equipment (MacBook Pro & iPhone zur privaten/beruflichen Nutzung). - Zugang zu Lernplattformen wie Pluralsight und O’Reilly Online Learning. - Ergonomisches Homeoffice-Setup nach Wunsch. - Wellpass-Membership Zuschuss für sportliche Aktivitäten wie Schwimmbäder, Fitness-Studios uvm. - 30 Tage Urlaub.
• Administration and support of service operations • Setting up build and code delivery processes (CI/CD) • Automation and optimization of routine tasks


