Job Closed
This listing is no longer active.
Passionate music fans. Innovative tech pros. Perfect harmony. Join our band.
Senior Site Reliability Engineer
Location
New York
Posted
134 days ago
Salary
$164.4K - $234.9K / year
Seniority
Senior
Job Description
Senior Site Reliability Engineer
Spotify
• Own fleet reliability. Lead the reliability, security, and scalability strategy for Portal’s SaaS infrastructure, including the runtime environments that power our platform and LLM-driven agent workflows. Define SLOs, drive capacity planning, and ensure our systems meet the demands of a rapidly growing product. • Architect for the agentic era. Design and evolve infrastructure on GCP and AWS using Terraform and infrastructure-from-code patterns. Shape how we structure environments for non-deterministic AI workloads — including sandboxing, resource isolation, cost governance, and security boundaries. • Drive operational excellence. Evolve our incident management, on-call, and postmortem practices. Leverage AI assistants to accelerate root cause analysis and build increasingly self-healing capabilities into our production systems. • Lead fullstack reliability. Operate across a modern web stack (TypeScript, React, Python). While not frontend-heavy, you’ll diagnose and resolve issues across the stack and drive reliability improvements end-to-end. • Mentor and multiply. Raise the reliability IQ of the broader engineering team. Establish SRE best practices, conduct production-readiness reviews, and mentor engineers on operational thinking. • Shape the roadmap. Partner with engineering and product leadership to evolve our infrastructure in step with generative AI features. Translate operational insights into strategic input on the product roadmap.
Job Requirements
- 5+ years of hands-on experience operating cloud infrastructure (GCP and/or AWS), using Terraform and Kubernetes to run production systems at scale.
- practical experience — or a strong demonstrated interest — in operating LLM-based systems, RAG pipelines, or agentic workloads, and understand the reliability challenges of non-deterministic systems.
- think in distributed systems first principles — consistency, availability, partition tolerance — and translate that thinking into pragmatic infrastructure decisions.
- proficient in at least one modern language (TypeScript, Java, Go, or Python) and comfortable navigating large, heterogeneous codebases, including environments where AI-generated PRs are common.
- build automation and improve systems so that whole categories of operational issues disappear over time.
- communicate complex infrastructure trade-offs clearly to both technical and non-technical stakeholders, and write postmortems that lead to meaningful change.
Benefits
- health insurance
- six-month paid parental leave
- 401(k) retirement plan
- monthly meal allowance
- 23 paid days off
- paid flexible holidays
- paid sick leave
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Lead design, development, and optimization of CI/CD pipelines using industry-leading tools • Implement robust Infrastructure as Code (IaC) for scalable, secure on-prem and cloud environments • Drive automation of provisioning, deployment, testing, and monitoring leveraging scripting languages (Python, Bash, golang) • Develop and manage containerized applications and orchestrate workloads with Docker and Kubernetes • Utilize cloud platforms (AWS, Azure) to optimize infrastructure performance, resilience, and cost • Implement and support continuous monitoring, logging, and alerting utilizing tools such as Prometheus, Grafana, and OpenTelemetry (OTEL) • Apply configuration management solutions (Ansible or other - Chef, Puppet, or SaltStack) • Champion DevSecOps practices by integrating security into all automated pipelines and infrastructure • Lead troubleshooting and resolution of complex deployment, infrastructure, and system issues • Collaborate cross-functionally with software engineering, QA, product management, security, and IT operations throughout agile, cross-disciplinary teams • Establish and maintain clear, comprehensive documentation including procedures, runbooks, technical solutions to support team knowledge and operational excellence • Mentor, coach, and provide technical leadership to DevOps engineers and advocate for best practices across engineering teams
• Design and implement infrastructure automation using IaC tools. • Build and maintain CI/CD pipelines. • Manage and optimize GKE clusters and containerized workloads. • Develop internal developer platforms and self-service tools. • Implement monitoring, logging, and alerting for platform reliability. • Collaborate with engineering and operations teams to improve workflows. • Ensure platform security, compliance, and scalability via automation. • Document architecture, processes, and tooling.
• Acompañar a los squads en la adopción de prácticas SRE • Prototipar soluciones técnicas innovadoras • Acelerar la adopción de estándares de calidad en toda la organización
• Collaborate with internal and customer development teams to enhance automation tools used in test engineering and software development • Configure, maintain, and monitor CI/CD pipeline automation within virtualization infrastructures • Contribute to DevOps libraries that support automation for our D4T product and internal DevSecOps processes • Engage with a variety of tools, data sources, and infrastructure technologies in support of new software projects • Assist in implementing security stages for our systems • Oversee the development and management of our Artifactory asset management system • Support and manage software build artifacts, including libraries, installers, Docker images, and NuGet packages • Develop and maintain our MLOps infrastructure for running ML/AI models and supporting data analytics pipelines • Set up and maintain Jenkins and Ansible servers • Collaborate with IT to address issues and create solutions, focusing on both on-premises and cloud-based services • Work directly with customers and their IT departments to implement DevOps for test infrastructure, which may include on-site services • Help design architecture requirements for customer collaboration, whether cloud-based or on-premises solutions are required • Participate in identifying future technology roadmaps and new business opportunities • Continuously invest in professional growth and stay current with modern technologies • Identify new possibilities for dashboards and data analytics enhancements




