Your app, Enterprise Ready.
Site Reliability Engineer
Location
United States
Posted
157 days ago
Salary
$175K - $275K / year
Seniority
Senior
Job Description
Site Reliability Engineer
WorkOS
• Design and evolve the systems, tooling, and processes that improve the reliability and performance of WorkOS • Collaborate with product and infrastructure teams to ensure services are production-ready, observable, and resilient to failure • Define and measure SLIs/SLOs to guide reliability improvements • Write and optimize backend systems (in TypeScript) with a focus on performance, maintainability, and graceful degradation • Improve our incident response process, lead postmortems, and drive follow-through on reliability risks • Develop internal tools and automations that make it easier to operate and scale our systems • Participate in our on-call rotation—responding to, resolving, and learning from production incidents • Contribute to design and architecture discussions with a focus on operability and long-term sustainability • Document systems, share learnings, and help grow a reliability-minded engineering culture
Job Requirements
- Experience operating and scaling production systems in cloud environments (we use AWS)
- Familiarity with service reliability concepts—monitoring, alerting, incident response, and root cause analysis
- Comfort working across infrastructure layers (e.g. compute, networking, storage, observability tooling)
- Strong debugging and systems thinking skills—you can follow problems across services and layers
- Ability to work independently, take ownership, and drive projects from problem discovery through resolution
- Nice to have*
- Familiarity with Kubernetes or similar orchestration systems
- Exposure to observability stacks (e.g. Prometheus, Grafana, Datadog, OpenTelemetry)
- Exposure to TypeScript or interest in working in a TypeScript-based codebase
Benefits
- Competitive pay
- Substantial equity grants
- Healthcare insurance (Medical, Dental and Vision) for you and your family
- 401k matching
- Wellness and fitness monthly allowances
- PTO + paid holidays + unlimited sick leave
- Autonomy and flexibility with remote work
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Network DevOps Engineer, RDMA Fabric Automation
VultrVultr is on a mission to make high-performance cloud computing easy to use, affordable, and locally accessible.
• Automate deployment and operations of large-scale RDMA (RoCEv2) Ethernet fabrics across Vultr data centers. • Build Ansible and Python-based frameworks to provision, validate, and remediate underlay and overlay networks. • Integrate network automation with Vultr’s source-of-truth systems (NetBox, OpsMill) for intent-driven configuration and validation. • Develop telemetry ingestion and correlation pipelines (gNMI, Prometheus, Kafka, custom collectors) for real-time network health and performance metrics. • Collaborate with platform, orchestration, and product engineering teams to optimize RDMA performance, PFC/ECN behavior, and path symmetry across fabrics. • Implement CI/CD workflows for network configuration changes — validation, pre-checks, and rollbacks. • Investigate complex network behaviors across layers — flow hashing, congestion domains, ECMP, and overlay interactions. • Contribute to the design of next-generation GPU and AI interconnect fabrics, ensuring seamless integration into Vultr’s global network architecture.
Senior Site Reliability Engineer, Core Cloud Engineering
VultrVultr is on a mission to make high-performance cloud computing easy to use, affordable, and locally accessible.
• Operate and scale Vultr’s control plane, ensuring availability, correctness, and performance across global datacenters. • Design, implement, and maintain automation to manage hypervisor fleets (KVM, QEMU, libvirt) and supporting infrastructure at scale. • Develop tooling and automation for Open vSwitch (OVS), BGP routing, and other networking components to ensure resilient and self-healing network operations. • Continuously analyze and improve system performance across compute, storage, and network layers, with an emphasis on reducing toil and eliminating single points of failure. • Implement advanced monitoring, logging, and tracing solutions (Grafana, Sentry, SumoLogic) while leading incident response to minimize impact and drive postmortem culture. • Maintain and evolve infrastructure pipelines (GitLab CI/CD, Puppet) to enable safe, fast, and reliable changes to both control plane and hypervisor infrastructure. • Work closely with Software Engineers, Network Engineers, and Product teams to align platform reliability with business and user needs. • Produce clear technical documentation for runbooks, operational procedures, and automation frameworks to improve team efficiency and reliability standards. • Coach and mentor team members in best practices for site reliability, incident handling, automation, and low-level Linux systems debugging.
DevOps Engineer
VultrVultr is on a mission to make high-performance cloud computing easy to use, affordable, and locally accessible.
• Build and maintain production-grade automation using Ansible, Terraform, and Go • Develop and enhance core services behind VKE, VLB (HAProxy), VCR (Container Registry), and Inference • Engage deeply with Kubernetes internals (scheduler, kubelet, controllers, CRDs) • Design, implement, and improve container runtime integrations (containerd, runc, OCI) • Architect CI/CD improvements and deployment pipelines for large-scale systems • Troubleshoot complex issues across networking, load balancing, containers, and distributed systems • Contribute directly to Vultr’s open-source ecosystem (terraform provider, crossplane, vultr-cli, govultr) • Improve overall reliability, observability, and operability of cloud-native services • Collaborate with product, platform, and infrastructure teams on feature delivery
• Design, build, and maintain scalable and resilient infrastructure on Microsoft Azure to support production SaaS workloads • Define and track service level objectives (SLOs), service level indicators (SLIs), and error budgets to drive reliability decisions • Build and maintain comprehensive monitoring, alerting, and observability systems to ensure early detection of issues • Develop and maintain CI/CD pipelines using GitHub Actions to enable safe, rapid, and repeatable deployments • Lead incident response and on-call rotations, conduct blameless post-incident reviews, and drive follow-up action items to completion • Automate operational tasks and eliminate toil through scripting, infrastructure-as-code, and self-healing systems • Manage and optimize Azure Kubernetes Service (AKS) clusters, container orchestration, and related networking and storage configurations • Collaborate with software engineering teams to embed reliability into application architecture, including capacity planning, load testing, and chaos engineering • Maintain and improve infrastructure-as-code using tools such as Terraform, Bicep, or ARM templates • Partner cross-functionally with Product, Support, and Quality to reduce friction and accelerate delivery

