Vultr is on a mission to make high-performance cloud computing easy to use, affordable, and locally accessible.
Senior Site Reliability Engineer, Core Cloud Engineering
Location
United States
Posted
98 days ago
Salary
$120K - $130K / year
Seniority
Senior
Job Description
Senior Site Reliability Engineer, Core Cloud Engineering
Vultr
• Operate and scale Vultr’s control plane, ensuring availability, correctness, and performance across global datacenters. • Design, implement, and maintain automation to manage hypervisor fleets (KVM, QEMU, libvirt) and supporting infrastructure at scale. • Develop tooling and automation for Open vSwitch (OVS), BGP routing, and other networking components to ensure resilient and self-healing network operations. • Continuously analyze and improve system performance across compute, storage, and network layers, with an emphasis on reducing toil and eliminating single points of failure. • Implement advanced monitoring, logging, and tracing solutions (Grafana, Sentry, SumoLogic) while leading incident response to minimize impact and drive postmortem culture. • Maintain and evolve infrastructure pipelines (GitLab CI/CD, Puppet) to enable safe, fast, and reliable changes to both control plane and hypervisor infrastructure. • Work closely with Software Engineers, Network Engineers, and Product teams to align platform reliability with business and user needs. • Produce clear technical documentation for runbooks, operational procedures, and automation frameworks to improve team efficiency and reliability standards. • Coach and mentor team members in best practices for site reliability, incident handling, automation, and low-level Linux systems debugging.
Job Requirements
- Proficiency in PHP with strong scripting and automation skills.
- Experience running large-scale distributed systems and control plane infrastructure in production.
- Strong background in hypervisor technologies (libvirt, QEMU, KVM) and Linux systems administration.
- Expertise in networking protocols and tools, particularly BGP and Open vSwitch (OVS), with automation experience.
- Deep knowledge of observability and monitoring frameworks (Grafana, Sentry, SumoLogic) and incident management.
- Advanced troubleshooting skills across compute, networking, and storage subsystems.
- Experience building and maintaining CI/CD pipelines (GitLab) and configuration management (Puppet).
- Familiarity with MySQL or similar databases, with an understanding of operational considerations for reliability and scale.
- Strong problem-solving abilities and the drive to tackle complex, low-level reliability challenges.
- Effective cross-team communication and collaboration skills.
- A commitment to continuous improvement and fostering a culture of operational excellence.
Benefits
- Excellent Medical Benefits w/ 100% company paid premiums for employee only plan + 100% company paid dental & vision premiums
- 401(k) plan that matches 100% up to 4% with immediate vesting
- Professional Development Reimbursement of $2,500 each year
- 11 Holidays + Paid Time Off Accrual + Rollover Plan
- Increased PTO at 3 year & 10 year anniversary + 1 month paid sabbatical every 5 years + Anniversary Bonus each year
- $500 first year remote office setup + $400 each following year for new equipment
- Internet reimbursement up to $75 per month
- Gym membership reimbursement up to $50 per month
- Company paid Wellable subscription
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Design, build, and maintain scalable and resilient infrastructure on Microsoft Azure to support production SaaS workloads • Define and track service level objectives (SLOs), service level indicators (SLIs), and error budgets to drive reliability decisions • Build and maintain comprehensive monitoring, alerting, and observability systems to ensure early detection of issues • Develop and maintain CI/CD pipelines using GitHub Actions to enable safe, rapid, and repeatable deployments • Lead incident response and on-call rotations, conduct blameless post-incident reviews, and drive follow-up action items to completion • Automate operational tasks and eliminate toil through scripting, infrastructure-as-code, and self-healing systems • Manage and optimize Azure Kubernetes Service (AKS) clusters, container orchestration, and related networking and storage configurations • Collaborate with software engineering teams to embed reliability into application architecture, including capacity planning, load testing, and chaos engineering • Maintain and improve infrastructure-as-code using tools such as Terraform, Bicep, or ARM templates • Partner cross-functionally with Product, Support, and Quality to reduce friction and accelerate delivery
Senior DevOps – Platform Engineer, Harness
XebiaCreating Digital Leaders. Digital Transformation Consultancy Services and Solutions
• Own and evolve the Harness platform while enabling fast, safe, and reliable cloud-native deployments across AWS, Azure, and GCP environments • Design and maintain Harness CI/CD pipelines for Kubernetes, ECS, Serverless, and VM workloads • Implement modern deployment strategies including Canary and Blue-Green releases • Build reusable pipeline templates and delivery workflows • Standardize infrastructure provisioning using Terraform and Helm / Kustomize • Embed Security, quality gates, and automated testing into CI/CD pipelines • Integrate Observability tooling and support platform reliability • Onboard and enable engineering teams on platform capabilities
• Architect and operate multi-region deployments across AWS, GCP, or Azure • Build and maintain high-throughput telemetry ingestion pipelines • Design autoscaling and failover strategies for mission-critical services • Own observability systems including Prometheus, Grafana, and distributed tracing • Improve MTTR and operational readiness processes • Manage CI/CD pipelines, GitOps workflows, and automated deployments • Collaborate with backend teams on API performance and infrastructure reliability • Harden infrastructure for security, compliance, and tenant isolation • Drive long-term infrastructure roadmap and architectural direction
Senior DevOps, Platform Engineer
XebiaCreating Digital Leaders. Digital Transformation Consultancy Services and Solutions
• own and evolve the Harness platform • enable fast, safe, and reliable cloud-native deployments across AWS, Azure, and GCP environments • design and maintain Harness CI/CD pipelines for Kubernetes, ECS, Serverless, and VM workloads • implement modern deployment strategies including Canary and Blue-Green releases • build reusable pipeline templates and delivery workflows • standardize infrastructure provisioning using Terraform and Helm / Kustomize • embed Security, quality gates, and automated testing into CI/CD pipelines • integrate Observability tooling and support platform reliability • onboard and enable engineering teams on platform capabilities



