Job Closed

This listing is no longer active.

RemoteWoman logo
RemoteWoman

Work remotely at trusted companies.

Site Reliability Engineer

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 1-10H1B No SponsorCompany SiteLinkedIn

Location

France

Posted

145 days ago

Salary

0

Seniority

Senior

Job Description

Site Reliability Engineer

RemoteWoman

• Refine Monitoring and Observability: Enhance system monitoring with tools like Prometheus, Grafana, and ELK Stack, ensuring visibility and alignment with business objectives. • Automate Deployments and Workflows: Transition manual processes to automated solutions using IaC tools (e.g., Terraform, Ansible) to streamline deployments and improve operational efficiency. • Optimize CI/CD Pipelines: Improve pipeline architecture for fast, reliable releases, ensuring scalability and resilience to handle high volumes of changes. • Cloud Infrastructure Management: Help scale cloud-based systems on platforms like AWS, GCP, and Azure while minimizing technical debt and operational complexity. • Incident Response and Post-Mortem: Support incident management and lead post-mortem analysis, ensuring continuous improvement and knowledge sharing. • Collaborate with Cross-Functional Teams: Work closely with engineering and product teams to integrate reliability practices into the development lifecycle and prioritize reliability efforts. • Drive Technical Innovation: Introduce and champion new tools, technologies, and practices that improve system reliability, performance, and scalability.

Job Requirements

  • DevOps, Cloud Operations, or SRE Expertise: A solid understanding of DevOps, Cloud Operations, or SRE principles, with a focus on reliability and scalability.
  • Advanced Linux Internals Expertise: Hands-on experience with Linux systems, including performance tuning, kernel configurations, and troubleshooting.
  • Programming Languages: Proficiency in programming languages such as Go (preferred) or Python, with a focus on building tools and automating processes.
  • Scripting Skills: Strong skills in scripting languages like Python, Bash, or Go to automate workflows, streamline tasks, and manage infrastructure.
  • Cloud Infrastructure Knowledge: Extensive experience with cloud platforms like AWS, GCP, and Azure, along with expertise in monitoring/logging frameworks and CI/CD pipelines.
  • Containerization and Orchestration: Hands-on experience with Docker, Kubernetes, and other containerization technologies for building and deploying scalable applications is a nice to have.
  • Problem-Solving and Collaboration: Strong problem-solving skills, system design experience, and the ability to collaborate effectively across teams.

Benefits

  • Flexible PTO
  • Comprehensive healthcare coverage (UK, France, Spain)
  • Company stock options
  • Professional development budget
  • Office equipment budget
  • Wellness budget
  • Annual team gatherings
  • Internet reimbursement
  • Inclusive parental leave
  • Remote work travel program

Related Categories

Related Job Pages

More DevOps Engineer Jobs

NeuroFlow logo

Senior DevOps Engineer, 6 month contract

NeuroFlow

Promoting behavioral health access and engagement in all care settings

DevOps Engineer145 days ago
OtherRemoteTeam 51-200Since 2017

• Develop, refactor, and maintain Terraform modules supporting our AWS infrastructure • Manage container deployments and service updates in Elastic Container Service (ECS) • Support small-scope feature requests and infrastructure improvements in AWS • Contribute to CI/CD pipeline development and performance tuning • Collaborate closely with the Platform Engineering team to propose and implement solutions

United States
Job Closed
Platform.sh logo

Site Reliability Engineer

Platform.sh

The PaaS that gives development teams control and peace of mind to deliver applications faster, at scale.

DevOps Engineer145 days ago
Full TimeRemoteTeam 201-500Since 2015H1B No Sponsor

• Refine Monitoring and Observability: Enhance system monitoring with tools like Prometheus, Grafana, and ELK Stack, ensuring visibility and alignment with business objectives. • Automate Deployments and Workflows: Transition manual processes to automated solutions using IaC tools (e.g., Terraform, Ansible) to streamline deployments and improve operational efficiency. • Optimize CI/CD Pipelines: Improve pipeline architecture for fast, reliable releases, ensuring scalability and resilience to handle high volumes of changes. • Cloud Infrastructure Management: Help scale cloud-based systems on platforms like AWS, GCP, and Azure while minimizing technical debt and operational complexity. • Incident Response and Post-Mortem: Support incident management and lead post-mortem analysis, ensuring continuous improvement and knowledge sharing. • Collaborate with Cross-Functional Teams: Work closely with engineering and product teams to integrate reliability practices into the development lifecycle and prioritize reliability efforts. • Drive Technical Innovation: Introduce and champion new tools, technologies, and practices that improve system reliability, performance, and scalability.

Germany
Job Closed
Full TimeRemoteTeam 10,001+Since 1978H1B No Sponsor

• Support and maintain infrastructure automation using Terraform or similar IaC tools • Develop and update CI/CD pipelines for development teams • Work with Kubernetes environments, including GKE cluster maintenance under senior guidance • Participate in building internal developer tools and platform components • Implement basic monitoring, logging, and alerting solutions • Troubleshoot platform issues and improve team workflows • Contribute to platform documentation and knowledge sharing

Ukraine
Job Closed
Open Systems logo

Senior Site Reliability Engineer

Open Systems

Managed SASE solutions that securely connect hybrid IT environments.

DevOps Engineer145 days ago
Full TimeRemoteTeam 201-500Since 1990

• Building Operational Automation: Design, build, and evolve the automation framework and tooling that powers the MC platform — primarily in Golang — with a strong focus on maintainability, scalability, and reliability. • Developing Self-Service APIs: Build and maintain the service APIs and self-service operational tooling of the MC platform that enable customers and teams to safely and efficiently operate services in production without manual intervention. • Applying Site Reliability Engineering (SRE) Principles: Define, implement, and continuously improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and SLA measurements so that reliability is measurable and actionable across the MC platform and the services it operates. • Owning Reliability and Operations Initiatives: Take ownership of reliability, automation, and Mission Control projects, driving them independently from problem identification through implementation and long-term operation. • Collaborating with AI Engineering: Work closely with our AI team and tooling, integrating AI-assisted capabilities into our automation and operational workflows. • Incident Response and Learning: Participate in incident response and the on-call rotation, leading root cause analysis and driving sustainable corrective and preventive actions. • Lead midsize to large automation and reliability initiatives, from early concept and design through production deployment and ongoing operations.

Switzerland