Flip logo
Flip

Empower Every Employee!

Senior Site Reliability Engineer

Location

Europe

Posted

4 days ago

Salary

0

Seniority

Senior

Job Description

Senior Site Reliability Engineer

Flip

• Co-owner of the architecture: Help drive the architecture and evolution of our cloud infrastructure on Azure and our Kubernetes clusters—designed for high throughput and maximum availability—to support Flip’s rapid global growth. • Drive the resilience strategy: Define our approach to global scaling, zero-downtime deployments, rollback mechanisms, and disaster recovery, ensuring the platform remains available around the clock. • Evolve our observability stack: Optimize our LGTM stack (Loki, Grafana, Tempo, Mimir) into a foundation our engineers can rely on. • Improve our IaC platform: Remove toil at the source and make our infrastructure a true self-service for engineering teams. • Lead during incidents: Take a leading role in major platform incidents, conduct factual post-incident analyses (blameless post-mortems), and turn findings into lasting improvements. • Mentor within the squad: Coach team members, lead RFCs and design reviews, and help engineers grow into stronger SREs. • Shape our roadmap: Collaborate closely with your squad to define the direction of the platform.

Job Requirements

  • 5+ years of hands-on experience as a Site Reliability Engineer (SRE), Platform Engineer, DevOps Engineer, Infrastructure Engineer, Cloud Engineer, or Backend Engineer with a strong infrastructure focus.
  • Proven track record of building and operating high-throughput, highly available systems in production.
  • Deep, production-grade experience with Kubernetes on one of the major hyperscalers.
  • Strong experience with modern observability stacks (e.g., Prometheus, Mimir, VictoriaMetrics, Dash0, Loki, ELK) and a clear understanding of SLIs, SLOs, and error budgets.
  • Solid software development skills in Go (strongly preferred, as our IaC runs on Pulumi in Go) or Python.
  • Hands-on experience with Infrastructure as Code (Pulumi, OpenTofu, Terraform) and GitOps (e.g., Argo CD), plus CI/CD pipeline design.
  • Demonstrated ability to lead complex infrastructure initiatives from design to production—including writing RFCs and driving architectural decisions within your team.
  • Experience mentoring engineers and raising the technical level within a team.
  • Comfortable taking end-to-end responsibility for critical incidents and turning insights into sustainable technical improvements.
  • Strong communication skills and fluent English.
  • Willingness to participate in on-call rotations to ensure the reliability of our platform.

Benefits

  • Work mode: We are remote-first, giving you the flexibility to work from home. At the same time, we value the benefits of in-person collaboration. Depending on the role, you will occasionally join team events, workshops, or meetings at our offices in Berlin or Stuttgart—always with sufficient notice. The exact balance will be discussed transparently during your interview process.
  • Work–life balance: We don’t want you glued to your desk, so we cover the cost of your E-Gym/Wellpass membership and offer company bike leasing (JobRad).
  • Celebrate success: You’ll work with highly motivated and engaged people in a relaxed working atmosphere.
  • Have impact: You’ll actively shape Flip. You’ll be an enabler of the rapid growth of a young tech company and grow alongside your goals. Good vibes guaranteed.
  • Happy to be a Flipster: Look forward to regular team events and culture days that bring us together as Flipsters.
  • Work abroad: At Flip you can also work from other European countries — let’s talk about workation during the interview.

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Full TimeRemoteTeam 51-200Since 2018

• Co-own the architecture: Help drive the architecture and evolution of our cloud infrastructure on Azure and our Kubernetes clusters - designed for high throughput and highest availability - to support Flip's rapid growth across the globe. • Drive the resilience strategy: Define how we approach global scaling, zero-downtime deployments, rollback mechanisms and disaster recovery, and make sure the platform stays available around the clock. • Evolve our observability stack: Improve our LGTM stack (Loki, Grafana, Tempo, Mimir) into a foundation our engineers can trust. • Improve our IaC Platform: Eliminate toil at the source, and make our infrastructure truly self-service for engineering teams. • Lead in incidents: Take a leading role in platform-related major incidents, drive blameless post-mortems for the squad, and translate findings into systemic improvements. • Mentor within the squad: Coach teammates, run RFCs and design reviews inside the team, and help engineers grow into stronger SREs. • Shape our roadmap: Partner with your squad to define the platform's direction.

Germany
OpsMill logo

Product Reliability Engineer

OpsMill

Great infrastructure automation starts with great infrastructure data.

DevOps Engineer4 days ago
Full TimeRemoteTeam 11-50Since 2023

• Partner directly with customers and with our Solution Architecture/Customer Success teams on L2/L3 escalations—communicating findings, driving root-cause analysis, and resolving complex packaging, deployment, upgrade, and runtime issues across heterogeneous Kubernetes environments. • Drive issues to resolution by reproducing problems locally, isolating root causes, and coordinating fixes with engineering—then documenting learnings in crisp RCAs that become actionable improvements • Build and maintain diagnostics tooling including support bundles, health checks, environment validators, and "what changed?" helpers that make future troubleshooting 10x faster • Own the test automation infrastructure roadmap, improving CI stability, reducing flaky tests, and creating reproducible integration/e2e environments that catch issues before customers do • Establish and maintain performance baselines and regression tests that serve as actionable gates, helping teams catch scale and latency issues early • Improve installation and upgrade robustness by identifying recurring failure modes and eliminating them through product changes, automation, and guardrails • Write production-quality code in Python, Go, or Rust for internal tooling and product improvements that directly enhance reliability • Close the reliability feedback loop by systematically turning field issues into better tests, observability, documentation, and product defaults—measuring success through reduced time-to-resolution and fewer repeat incidents

United States
OpsMill logo

Product Reliability Engineer

OpsMill

Great infrastructure automation starts with great infrastructure data.

DevOps Engineer4 days ago
Full TimeRemoteTeam 11-50Since 2023

• Partner directly with customers and with our Solution Architecture/Customer Success teams on L2/L3 escalations—communicating findings, driving root-cause analysis, and resolving complex packaging, deployment, upgrade, and runtime issues across heterogeneous Kubernetes environments. • Drive issues to resolution by reproducing problems locally, isolating root causes, and coordinating fixes with engineering—then documenting learnings in crisp RCAs that become actionable improvements • Build and maintain diagnostics tooling including support bundles, health checks, environment validators, and "what changed?" helpers that make future troubleshooting 10x faster • Own the test automation infrastructure roadmap, improving CI stability, reducing flaky tests, and creating reproducible integration/e2e environments that catch issues before customers do • Establish and maintain performance baselines and regression tests that serve as actionable gates, helping teams catch scale and latency issues early • Improve installation and upgrade robustness by identifying recurring failure modes and eliminating them through product changes, automation, and guardrails • Write production-quality code in Python, Go, or Rust for internal tooling and product improvements that directly enhance reliability • Close the reliability feedback loop by systematically turning field issues into better tests, observability, documentation, and product defaults—measuring success through reduced time-to-resolution and fewer repeat incidents

Czechia
Sensor Tower logo

DevOps Engineer

Sensor Tower

Better data comes from real people

DevOps Engineer4 days ago
Full TimeRemoteTeam 201-500

• Collaborate closely with our engineering team to help them be more productive by improving development tools. • Maintain a strong connection to the product and remain comfortable engaging in direct feature development • You will take ownership of CI/CD stack, keeping it robust, observable, stable, and performant. • Research, implement, and develop tools to help developers write new code. • Ensure the CI/CD infrastructure and process is reliable and consistent.

Poland