Staff Site Reliability Engineer
Location
PST (UTC-8)
Posted
31 days ago
Salary
$200K - $270K / year
Seniority
Lead
No structured requirement data.
Job Description
Staff Site Reliability Engineer
Bluesky
Role Description We're looking for a Staff Site Reliability Engineer to help design, implement, and operate the infrastructure that powers Bluesky and atproto. This is a hands-on role for someone who has operated high-scale production systems, understands how distributed systems fail, and wants to build the operational foundation for an open social network. - Work across bare-metal systems, cloud services, data infrastructure, observability, incident response, capacity planning, and reliability engineering for systems serving millions of users. - Own reliability, availability, and operational excellence for our production systems, including observability, incident response, deployment, and rollback systems. - Improve production readiness for services, migrations, and infrastructure changes. - Develop software that pushes the state of the art in performance, automation, observability, and other areas. - Scale systems running on dense, latest-generation, bare-metal servers in our own colocation facilities. - Reduce toil through automation, tooling, and thoughtful engineering practices. - Partner with engineers across all our teams to help design services with strong operational characteristics. - Lead incident reviews and turn contributing factors into concrete engineering improvements as we practice continuous improvement. - Perform capacity planning and cost management across compute, storage, database, and networking workloads. - Manage various vendor relationships to ensure we can provide high quality services at a reasonable TCO. - Mentor engineers on reliability, operability, debugging, and distributed systems practices and help define a culture of operational excellence across the org. Qualifications - +10 years experience operating high-scale production systems, including bare metal. - Strong fundamentals in Linux, networking, storage, databases, and distributed systems. - Built and operated high-scale systems where correctness, latency, throughput, and availability were critical. - Can write production-quality software in Go. - Comfortable debugging across application code, operating systems, databases, networks, and hardware. - Experience with observability systems, alert design, incident response, capacity planning, Kubernetes, and production automation. - Like working on very small, fast-moving teams at a startup. - Have read the AT Protocol docs, feel aligned with the mission, and want to contribute! Requirements - Fully remote team with an overlap of working hours with PST required. - Willingness to travel to team meetups once every 3-4 months. - On-site component may be required for interviews. Benefits - Health, dental, and vision insurance. Additional Notes The anticipated base salary range for this position is $200,000 - $270,000 USD, excluding equity. Equity will be considered in the total compensation package. Final base salary for this role will be based on the individual's geographic location, as well as experience level, skill set, training, licenses, and certifications.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Government DevOps Engineer
HEREThe product of years of collaboration with the world’s largest financial institutions, HERE Enterprise Browser is the first and only browser that solves both enterprise security and workforce productivity. Built on Google Chromium, HERE streamlines workflow and improves employee experience.
Role Description HERE is seeking a Government focused DevOps Engineer to join our team! The primary responsibilities for this role will span CI/CD pipeline engineering and cloud operations, maintaining and improving our GitHub CI/CD pipelines, and supporting our AWS cloud infrastructure. In this role, you will grow your hands-on experience with real production build systems and cloud platforms while having the opportunity to work on practical projects that directly impact both our development velocity and operational reliability. You will play a vital role in ensuring our infrastructure complies with federal standards, directly supporting the delivery of our secure browser environment to public sector clients. We're actively evolving toward a cloud-agnostic, multi-cloud architecture and migrating to Kubernetes for container orchestration. While current AWS and ECS experience is essential, having exposure to Azure, GCP, and Kubernetes will position you well for our infrastructure roadmap. Responsibilities - CI/CD Pipeline Development: - Build, maintain, and optimize CI/CD pipelines for multi-platform builds (Windows, macOS, Linux). - Work with YAML configurations, pipeline stages, artifacts, and deployment workflows. - Integrate security and vulnerability scanning tools directly into the CI/CD pipeline to support automated compliance validations (DevSecOps). - Cloud Infrastructure Operations: - Help maintain and improve AWS infrastructure including ECS/Fargate deployments, RDS databases, Route53 DNS, VPC networking, and IAM policies. - Support multi-tenant, multi-region, and highly isolated or public-sector specific cloud architectures (e.g., AWS GovCloud deployments). - Container & Deployment Management: - Work with Docker containers, ECS task definitions, and ECR registries. - Deploy and manage containerized Node.js applications in production environments. - Assist in the implementation of hardened container base-images aligned with federal or highly-regulated industry security benchmarks. - Release Management: - Help manage release processes including version promotion, release channels (canary, beta, stable), and automated deployment to staging and production environments. - Database Operations: - Support PostgreSQL on AWS RDS—backups, SSH tunneling through bastion hosts, read-only user management, and database configuration for multi-tenant environments. - Automation & Scripting: - Write and maintain automation scripts in Bash, PowerShell, Python, and Node.js. - Build tools to improve infrastructure reliability and developer experience. - Internal Tools Support: - Help maintain web-based DevOps tools built with Express.js, React, and TypeScript—tools for cloud settings management, tenant provisioning, and deployment monitoring. Qualifications - Ideally 2 to 4 years of experience with the following core requirements: - GitLab CI/CD: Experience with GitLab CI/CD pipelines—YAML configuration, stages, jobs, artifacts, rules, dependencies. - Understanding of CI/CD best practices and pipeline optimization. - AWS Cloud Fundamentals: Production level experience with core AWS services—EC2, ECS/Fargate, RDS, Route53, VPC, IAM, Secrets Manager, CloudWatch. Comfortable navigating the AWS Console and CLI. - Multi-Platform Scripting: Solid scripting skills in Bash (Linux) and PowerShell (Windows). Ability to write maintainable automation scripts for both platforms. - Containerization: Hands-on Docker experience—building images, writing Dockerfiles, docker-compose, understanding container networking, and working with ECS/ECR. - Build Systems: Experience with build tools and package managers—npm/Node.js, .NET/NuGet, Python packaging. Understanding of dependency management and build artifacts. - Version Control: Strong Git fundamentals—branching strategies, merge requests, tagging. Experience with GitHub (or GitLab) workflows and code review practices. - Linux/Unix & Windows: Comfortable in both environments—SSH, file permissions, package managers, systemd, PowerShell. Understanding of cross-platform operational challenges. - Node.js/JavaScript: Comfortable reading and writing JavaScript/Node.js code. Experience with npm, package.json, and basic Express.js applications for tooling. - Functional knowledge of federal compliance frameworks like FedRAMP, NIST SP 800-53, DISA STIGs, or DoD Cloud SRG (IL4-IL6). - U.S. citizenship is mandatory; holding an active Secret clearance is preferred, or the ability to obtain one as required. - Ability to function effectively under stringent change control processes, regular auditing, and detailed documentation standards. Nice to Have - Kubernetes experience (EKS, GKE, AKS) or willingness to learn, we're migrating from ECS to K8s. - Multi-cloud experience (Azure, GCP) or cloud-agnostic architecture knowledge. - GitLab Runner administration and configuration. - AWS CDK or CloudFormation for Infrastructure as Code. - Terraform for multi-cloud infrastructure management. - TypeScript development experience. - PostgreSQL database administration and optimization. - .NET build systems and NuGet package management. - React or frontend framework experience. - Airflow or workflow orchestration tools. - Helm charts and Kubernetes manifest management. - Familiarity with FIPS 140-2/3 cryptographic compliance standards. - Hands-on experience with GitHub Actions administration and environment scaling. - Exposure to enterprise secret management tools like HashiCorp Vault or CyberArk. - Direct support of ATO (Authority to Operate) processes and eMASS documentation. Benefits - Generous Paid Time Off, Paid Holidays & Sick Time - Competitive & Comprehensive Health Insurance - Thoughtfully-Planned Paid Parental Leave - Financial Well-Being Plans (FSA) (401k) (Life Insurance) - Stock Options - Professional Development Courses - Employee Resource Groups Additional Perks - One Medical - Free Membership - Talkspace - Mental Health Therapy 24/7 - Team Lunches - Casual dress code - Commuter Benefits (NYC employees only) - Citibike (NYC employees only) Salary Range $145k - $185k This base salary range represents the low and high end salary range for this particular position; not all encompassing of the total compensation package. Actual salaries may vary depending upon but not limited to experience, special skill set, education and location. This range represents only one aspect of HERE’s total compensation package offered to employees. Other forms of compensation may be stock options, commissions, paid time off and other variable benefits.
DevOps & Cloud Security Manager
ultima millaLogistic Management System for E-commerce & Retail in Mexico. Raised +$7M USD from Y Combinator, FJLabs, & more.
• Liderarás la estrategia de infraestructura cloud y seguridad de la compañía en entornos GCP y AWS. • Definir y ejecutar la estrategia de seguridad cloud: zero trust, segmentación de red, mínimo privilegio y defense-in-depth en GCP y AWS. • Implementar controles preventivos y detectivos: WAF, DDoS mitigation, IDS/IPS, SIEM y gestión de vulnerabilidades (SAST, DAST, CVE tracking). • Integrar seguridad en el pipeline CI/CD (shift-left): secret scanning, image scanning y análisis de dependencias. • Liderar respuesta a incidentes, threat modeling y coordinación de ejercicios de red team / pentesting externo. • Diseñar y mantener infraestructura cloud en GCP (GKE, Cloud Run, Cloud SQL, Pub/Sub, IAM) y AWS (ECS/Fargate, EKS, RDS, SQS, VPC, IAM). • Gestionar IaC (Terraform), pipelines CI/CD, observabilidad (Prometheus, Grafana, DataDog) y confiabilidad (SLOs/SRE).
Senior Site Reliability Engineer, C#, .NET
ClimavisionWe're rebuilding climate technology from the ground up.
• Own production reliability for Climavision’s customer-facing platform and radar-derived weather data services across Azure, colocation, and edge Kubernetes environments. • Contribute to the definition and improvement of SLIs, SLOs, alerting standards, and operational metrics used to measure platform reliability. • Support and coordinate production incident response efforts, including troubleshooting, mitigation, communication, and postmortem analysis. • Diagnose and resolve complex production issues across application services, Kubernetes infrastructure, storage, and distributed systems. • Drive multi-replica and multi-cluster high availability across Climavision’s .NET services. • Improve reliability and operational maturity of production platform services, including observability, autoscaling, ingress, and distributed storage. • Partner with software engineering teams to improve production readiness, resiliency patterns, deployment safety, and operational visibility before services reach production. • Support and evolve Climavision’s observability platform, including metrics, logging, distributed tracing, dashboarding, and alerting.
Role Description We're looking for a Middle SRE to join our team and help us build and maintain reliable, observable, and scalable infrastructure. You'll work closely with developers, own reliability practices, and contribute to the team's DevOps culture. Qualifications - Linux — confident working in the terminal, troubleshooting, system basics - Docker — understanding of containers, images, and basic Docker operations - GCP (Google Cloud Platform) — hands-on experience required; our entire infrastructure runs on GCP - Kubernetes — solid hands-on experience; understanding of workloads, networking, and cluster operations - GitOps / FluxCD — experience with FluxCD or similar GitOps tooling is a strong advantage - Observability — Prometheus, Alertmanager, Grafana; understanding the difference between metrics, logs, and traces - Git / GitHub — comfortable with branching strategies, PRs, and Git-based workflows - Knows what SLI, SLO, and SLA mean. - Basic understanding of AI-related concepts: agents, skills, MCP (Model Context Protocol) Requirements - CI/CD Pipelines — experience writing or maintaining pipelines (GitHub Actions, GitLab CI, or similar) - Databases — familiarity with PostgreSQL, Couchbase, or other databases; understanding of basic operations and monitoring - Networking fundamentals — knows where load balancers, gateways, and DNS fit in a cloud architecture - Infrastructure as Code — Terraform or similar IaC tooling Responsibilities - Managing and improving Kubernetes-based infrastructure - Maintaining and evolving observability stack (metrics, logs, traces) - Writing and improving CI/CD pipelines - Supporting and improving GitOps workflows - Collaborating with developers on reliability, performance, and scalability topics


