Job Closed
This listing is no longer active.
Work remotely at trusted companies.
Site Reliability Engineer
Location
France
Posted
145 days ago
Salary
0
Seniority
Senior
Job Description
Site Reliability Engineer
RemoteWoman
• Refine Monitoring and Observability: Enhance system monitoring with tools like Prometheus, Grafana, and ELK Stack, ensuring visibility and alignment with business objectives. • Automate Deployments and Workflows: Transition manual processes to automated solutions using IaC tools (e.g., Terraform, Ansible) to streamline deployments and improve operational efficiency. • Optimize CI/CD Pipelines: Improve pipeline architecture for fast, reliable releases, ensuring scalability and resilience to handle high volumes of changes. • Cloud Infrastructure Management: Help scale cloud-based systems on platforms like AWS, GCP, and Azure while minimizing technical debt and operational complexity. • Incident Response and Post-Mortem: Support incident management and lead post-mortem analysis, ensuring continuous improvement and knowledge sharing. • Collaborate with Cross-Functional Teams: Work closely with engineering and product teams to integrate reliability practices into the development lifecycle and prioritize reliability efforts. • Drive Technical Innovation: Introduce and champion new tools, technologies, and practices that improve system reliability, performance, and scalability.
Job Requirements
- DevOps, Cloud Operations, or SRE Expertise: A solid understanding of DevOps, Cloud Operations, or SRE principles, with a focus on reliability and scalability.
- Advanced Linux Internals Expertise: Hands-on experience with Linux systems, including performance tuning, kernel configurations, and troubleshooting.
- Programming Languages: Proficiency in programming languages such as Go (preferred) or Python, with a focus on building tools and automating processes.
- Scripting Skills: Strong skills in scripting languages like Python, Bash, or Go to automate workflows, streamline tasks, and manage infrastructure.
- Cloud Infrastructure Knowledge: Extensive experience with cloud platforms like AWS, GCP, and Azure, along with expertise in monitoring/logging frameworks and CI/CD pipelines.
- Containerization and Orchestration: Hands-on experience with Docker, Kubernetes, and other containerization technologies for building and deploying scalable applications is a nice to have.
- Problem-Solving and Collaboration: Strong problem-solving skills, system design experience, and the ability to collaborate effectively across teams.
Benefits
- Flexible PTO
- Comprehensive healthcare coverage (UK, France, Spain)
- Company stock options
- Professional development budget
- Office equipment budget
- Wellness budget
- Annual team gatherings
- Internet reimbursement
- Inclusive parental leave
- Remote work travel program
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Senior DevOps Engineer, 6 month contract
NeuroFlowPromoting behavioral health access and engagement in all care settings
• Develop, refactor, and maintain Terraform modules supporting our AWS infrastructure • Manage container deployments and service updates in Elastic Container Service (ECS) • Support small-scope feature requests and infrastructure improvements in AWS • Contribute to CI/CD pipeline development and performance tuning • Collaborate closely with the Platform Engineering team to propose and implement solutions
Site Reliability Engineer
Platform.shThe PaaS that gives development teams control and peace of mind to deliver applications faster, at scale.
• Refine Monitoring and Observability: Enhance system monitoring with tools like Prometheus, Grafana, and ELK Stack, ensuring visibility and alignment with business objectives. • Automate Deployments and Workflows: Transition manual processes to automated solutions using IaC tools (e.g., Terraform, Ansible) to streamline deployments and improve operational efficiency. • Optimize CI/CD Pipelines: Improve pipeline architecture for fast, reliable releases, ensuring scalability and resilience to handle high volumes of changes. • Cloud Infrastructure Management: Help scale cloud-based systems on platforms like AWS, GCP, and Azure while minimizing technical debt and operational complexity. • Incident Response and Post-Mortem: Support incident management and lead post-mortem analysis, ensuring continuous improvement and knowledge sharing. • Collaborate with Cross-Functional Teams: Work closely with engineering and product teams to integrate reliability practices into the development lifecycle and prioritize reliability efforts. • Drive Technical Innovation: Introduce and champion new tools, technologies, and practices that improve system reliability, performance, and scalability.
• Support and maintain infrastructure automation using Terraform or similar IaC tools • Develop and update CI/CD pipelines for development teams • Work with Kubernetes environments, including GKE cluster maintenance under senior guidance • Participate in building internal developer tools and platform components • Implement basic monitoring, logging, and alerting solutions • Troubleshoot platform issues and improve team workflows • Contribute to platform documentation and knowledge sharing
Senior Site Reliability Engineer
Open SystemsManaged SASE solutions that securely connect hybrid IT environments.
• Building Operational Automation: Design, build, and evolve the automation framework and tooling that powers the MC platform — primarily in Golang — with a strong focus on maintainability, scalability, and reliability. • Developing Self-Service APIs: Build and maintain the service APIs and self-service operational tooling of the MC platform that enable customers and teams to safely and efficiently operate services in production without manual intervention. • Applying Site Reliability Engineering (SRE) Principles: Define, implement, and continuously improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and SLA measurements so that reliability is measurable and actionable across the MC platform and the services it operates. • Owning Reliability and Operations Initiatives: Take ownership of reliability, automation, and Mission Control projects, driving them independently from problem identification through implementation and long-term operation. • Collaborating with AI Engineering: Work closely with our AI team and tooling, integrating AI-assisted capabilities into our automation and operational workflows. • Incident Response and Learning: Participate in incident response and the on-call rotation, leading root cause analysis and driving sustainable corrective and preventive actions. • Lead midsize to large automation and reliability initiatives, from early concept and design through production deployment and ongoing operations.




