Job Closed
This listing is no longer active.
faster builds that keep engineers in flow
Site Reliability Engineer
Location
California
Posted
157 days ago
Salary
0
Seniority
Senior
Job Description
Site Reliability Engineer
EngFlow
• Design, build, and maintain cloud infrastructure for our distributed build acceleration platform • Automate everything: from deployment pipelines to monitoring and recovery • Manage scalability and reliability for high-throughput, low-latency systems • Implement and maintain observability: logging, metrics, tracing, and alerting • Work closely with product and engineering teams to embed reliability into every feature • Diagnose and resolve production incidents quickly, and feed learnings back into systems design • Optimize cost, performance, and resilience across multi-cloud environments
Job Requirements
- 4+ years in SRE, DevOps, or Production Engineering roles
- Experience managing Kubernetes in production
- Strong background in cloud infrastructure (GCP or AWS) and IaC (Terraform preferred)
- Solid knowledge of networking, security, and distributed systems
- Track record of improving system availability and developer productivity
- A knack for debugging complex, cross-system issues under pressure
Benefits
- comprehensive medical, dental, vision benefits
- 401k/pension
- parental leave
- generous vacation
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Senior DevOps Engineer
Meazure LearningOffering full-service assessment development, delivery, and proctoring solutions across the world
• Help define and elevate the user experience for learners and professionals around the world • Collaborate with talented, mission-driven colleagues across regions • Work in a culture that values trust, innovation, and transparency • Design, implement, and maintain scalable, secure, and reliable CI/CD pipelines • Manage and optimize cloud infrastructure (e.g., AWS, Azure) and container orchestration (e.g., Kubernetes) • Drive automation across infrastructure and development workflows • Build and maintain monitoring, alerting, and logging systems to ensure reliability and observability • Collaborate with Engineering, QA, and Security teams to deliver high-performing, compliant solutions • Troubleshoot complex system issues in staging and production environments • Guide and mentor junior engineers and contribute to DevOps best practices
• Implement and maintain Infrastructure as Code (IaC) using Terraform. • Develop, manage, and optimize Terraform modules. • Manage and optimize AWS Cloud infrastructure, ensuring performance, scalability, and availability requirements are met. • Assist in implementing and maintaining monitoring and logging systems (e.g., ELK, Grafana, Thruk). • Provide technical support for infrastructure, CI/CD pipelines, and deployment processes. • Troubleshoot and resolve infrastructure-related issues, ensuring smooth operations. • Document and maintain DevOps workflows, processes, and system configurations.
• Lead and mentor a team of DevOps engineers, fostering technical growth and collaboration • Define and drive the infrastructure roadmap aligned with company objectives • Architect and oversee cloud infrastructure design and implementation • Establish best practices, standards, and processes for infrastructure development and operations • Partner with Engineering, Research, and FDE to align infrastructure capabilities with business needs • Drive the evolution of Kubernetes clusters optimized for GPU workloads, Production SaaS hosting and varied enterprise deployment models • Champion GitOps practices using ArgoCD for continuous deployment • Establish infrastructure as code standards using Terraform • Define monitoring and observability strategy for distributed systems • Collaborate with ML engineers to optimize infrastructure for model training and serving • Own infrastructure reliability, performance, and security posture • Implement and maintain cost optimization strategies (FinOps) for cloud resources
• Design and implement cloud infrastructure from the ground up • Build and maintain Kubernetes clusters optimized for GPU workloads and ML applications, as well as Production SaaS hosting • Implement GitOps practices using ArgoCD for continuous deployment • Develop infrastructure as code using Terraform • Create and maintain CI/CD pipelines for infrastructure and application deployment • Implement monitoring and observability solutions for distributed systems • Automate infrastructure management with Python and Bash • Collaborate with ML engineers to optimize infrastructure for model training and serving • Implement and maintain cost optimization strategies (FinOps) for cloud resources • Monitor and optimize cloud spending, especially for GPU-intensive workloads



