Defining what it means to build and deliver the most extraordinary sports & entertainment experiences.The Crown is Yours
Database Reliability Engineer
Location
United States
Posted
12 hours ago
Salary
$112K - $140K / year
Seniority
Mid Level
Job Description
Database Reliability Engineer
DraftKings Inc.
• Drive database reliability across PostgreSQL, MySQL, MongoDB, Redis, ScyllaDB, Aerospike, and managed cloud services • Support high availability, replication, partitioning, storage, and connection management • Develop Kubernetes operators, infrastructure as code, GitOps workflows, and production tooling in Go or Python • Automate provisioning, failover, backups, schema migrations, and lifecycle management • Build self-healing, fault-tolerant infrastructure and internal tooling to reduce operational toil • Optimize database performance and cost across cloud and on-premises environments • Support capacity planning, resource efficiency, storage optimization, workload consolidation, and performance tuning • Partner with application engineering teams on schema reviews, migration strategies, query optimization, connection management, and zero-downtime deployments • Leverage AI for observability, anomaly detection, root cause analysis, documentation, predictive insights, and code evaluation
Job Requirements
- At least 2 years of experience in Database Reliability Engineering, Database Platform Engineering, Site Reliability Engineering, or related infrastructure engineering with a strong database focus
- Experience supporting at least one major relational database, preferably Aurora MySQL
- Operational experience with PostgreSQL, MongoDB, ScyllaDB, Aurora, or other managed cloud database services
- Experience building and operating stateful Kubernetes workloads using StatefulSets, Persistent Volumes, database operators, Terraform, Pulumi, FluxCD, ArgoCD, GKE, or EKS
- Hands-on software development experience using Go or Python for automation, platform tooling, Kubernetes controllers, APIs, and infrastructure
- Experience with observability, monitoring, service level objectives, capacity planning, performance optimization, and self-service engineering solutions
- Practical experience using AI tools such as Claude, GitHub Copilot, Cursor, MCP, or similar technologies
- Sound engineering judgment to validate AI-generated outputs for reliability and security
- May be required to obtain a gaming license issued by the appropriate state agency as a condition of employment
Benefits
- Bonus
- Equity
- Benefits as applicable
- Guidance through the gaming-license process if required
- Remote work option within the US
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Senior DevOps Engineer, Infrastructure – Reliability
Worth AIAI Data-Driven Credit Score For Every Business 💸✨
• Build scalable Infrastructure-as-Code patterns with Terraform to standardize cloud provisioning and reduce configuration drift • Own and evolve the Kubernetes platform, ensuring workloads are secure, scalable, and resilient • Optimize CI/CD pipelines to improve deployment frequency, reduce lead time, and increase release confidence • Design and enforce secure networking, IAM, and secrets management strategies • Improve observability through metrics, logs, and tracing using tools such as DataDog • Optimize cloud costs through rightsizing, autoscaling, and architectural improvements • Implement disaster recovery, backup, and multi-region resilience initiatives • Refactor brittle or manual infrastructure into automated, testable, reproducible systems • Introduce infrastructure tooling or architectural changes and drive adoption through documentation, workshops, and hands-on support • Partner with engineering teams to reduce friction in CI/CD, deployments, and cloud environments • Communicate technical trade-offs across engineering and product stakeholders • Maintain or exceed SLO/SLA targets, reduce incident frequency and duration, increase infrastructure automation, and improve cloud cost efficiency
Role Description You are an innovative Site Reliability engineer with experience in managing large-scale SaaS deployments. You are comfortable in architecting deployment solutions that accelerate software deployment at scale, with a mindset towards reliability, observability and fault tolerance. You enjoy staying on top of latest technologies, and are always looking for ways to improve our deployment and tooling. You are comfortable presenting your solutions to internal teams, defining tasks, driving consensus and executing your plans. You are not afraid to be creative with technologies, and enjoy managing complex operations. Your Impact: - Work in an agile, "startup-style" environment with minimal bureaucracy. - Be part of a small team driving the next evolution of security at Cisco. - Work alongside some of the brightest minds in the industry in taking Cisco's Mini-Me solution to market. - Implement the infrastructure needed to deliver our Management Plane frontend to AWS for large-scale availability. - Create solutions to manage scale, performance and disaster recovery for high-use, frequently-updated software in an observable, reliable and fault-tolerant deployment. - Work closely with the DevOps engineers on a wide range of continuous improvement initiatives to help improve our infrastructure and customer experience. Qualifications - BA/BS Degree and 13-16 years of experience as a Site Reliability Engineer. - 10+ years in cloud computing and delivering SaaS applications. - Prior experience managing large-scale SaaS deployments to public cloud: security, load-balancing, logging, monitoring, tuning, redundancies, and deployment automation. - Strong background in Kubernetes and other containerization technologies. Requirements - Strong foundation in CI/CD tools and automation processes (e.g., Terraform, ArgoCD). - Working as part of a remote or distributed team. - Fluency in Linux environments. Company Description At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint. Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere. We are Cisco, and our power starts with you.
Role Description This is a remote position. - L1 support (AWS infrastructure/app health Monitoring - No on-prem infrastructure) - Incident Response and Operational support - Acknowledge / Respond to alerts from tools like Splunk, Datadog, Incident tickets etc. - Follow-up runbooks/SOPs - Escalate unresolved issues to L2/L3 team - Track MTTA/MTTR - Basic AWS admin functions (EC2, S3, IDM, Access Control, Backups, Scheduled jobs, restart VMs/services etc.) - Develop and maintain CI/CD pipelines and infrastructure as code (IaC) to support Plume Open Sync Cloud software and infrastructure updates enabling Cellular Backup (CBU) - Troubleshoot and resolve infrastructure and deployment issues - Implement cloud and security best practices to improve reliability and scalability - Collaborate with development and operations teams to optimize deployments - Automate infrastructure provisioning and configuration management - Enhance monitoring and observability using Prometheus, Grafana, and ELK - Document processes and assist in knowledge sharing - Shift handover reporting Company Description
• Design, build, and maintain secure and scalable cloud infrastructure within Google Cloud Platform • Support deployment, monitoring, reliability, and performance of cloud-based applications and services • Develop and improve infrastructure automation and deployment processes • Build and maintain CI/CD pipelines • Partner with engineering teams to streamline application releases and cloud operations • Monitor production environments and assist with troubleshooting, incident response, and root cause analysis • Improve cloud security, access controls, logging, monitoring, and operational standards • Identify opportunities to improve infrastructure reliability, scalability, and cost efficiency • Create and maintain technical documentation, runbooks, and operational procedures • Collaborate with distributed technical teams across locations and time zones



