Empower Every Employee!
Senior Site Reliability Engineer – m/f/d
Location
Germany
Posted
21 hours ago
Salary
0
Seniority
Senior
Job Description
Senior Site Reliability Engineer – m/f/d
Flip
• Co-own the architecture: Help drive the architecture and evolution of our cloud infrastructure on Azure and our Kubernetes clusters - designed for high throughput and highest availability - to support Flip's rapid growth across the globe. • Drive the resilience strategy: Define how we approach global scaling, zero-downtime deployments, rollback mechanisms and disaster recovery, and make sure the platform stays available around the clock. • Evolve our observability stack: Improve our LGTM stack (Loki, Grafana, Tempo, Mimir) into a foundation our engineers can trust. • Improve our IaC Platform: Eliminate toil at the source, and make our infrastructure truly self-service for engineering teams. • Lead in incidents: Take a leading role in platform-related major incidents, drive blameless post-mortems for the squad, and translate findings into systemic improvements. • Mentor within the squad: Coach teammates, run RFCs and design reviews inside the team, and help engineers grow into stronger SREs. • Shape our roadmap: Partner with your squad to define the platform's direction.
Job Requirements
- 5+ years of hands-on experience as a Site Reliability Engineer (SRE), Platform Engineer, DevOps Engineer, Infrastructure Engineer, Cloud Engineer, or Backend Engineer with a strong infrastructure focus.
- Proven track record building and operating **high-throughput, highly available systems** in production.
- Deep, production-level experience with **Kubernetes** on any Hyperscaler.
- Strong experience with modern observability stacks (e.g. Prometheus, Mimir, VictoriaMetrics, Dash0, Loki, ELK) and a clear point of view on SLIs, SLOs and error budgets.
- Solid software development skills in **Go** (strongly preferred, since our IaC runs on Pulumi in Go) or Python.
- Hands-on experience with Infrastructure as Code (Pulumi, OpenTofu, Terraform) and **GitOps** (e.g. ArgoCD) **+ CI/CD pipeline design**.
- Demonstrated ability to lead complex infrastructure initiatives from design to production - including writing RFCs and driving architecture decisions within your team.
- Experience mentoring engineers and raising the technical bar within a team.
- Comfortable owning major incidents end-to-end and turning learnings into systemic change.
- Strong communication skills and business-fluent English.
- Willingness to participate in on-call rotations to ensure the reliability of our platform.
Benefits
- Work mode: We’re remote-first, giving you flexibility to work from home. At the same time, we deeply value the power of in-person collaboration. Depending on the role, you’ll join occasional team events, workshops, or meetings in our Berlin or Stuttgart offices - always with plenty of notice. The exact balance will be discussed during your interview.
- Work-Life-Balance: We don't want you to grow roots to your desk chair. That's why we cover the costs of your E-Gym-Wellpass membership and offer job bike leasing.
- Celebrating success: Expect highly motivated and committed people in a relaxed working atmosphere.
- Be part of something bigger: You actively shape Flip in your role. Along the way, you are an enabler of the rapid growth process of a young tech company and grow towards your goals, fun is guaranteed.
- Happy to be a Flipster: Stay tuned for regular team events and culture days that bring us together as Flipsters.
- Working abroad: At Flip you can also work abroad in the European Union. Let's talk about remote work in the interview.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Product Reliability Engineer
OpsMillGreat infrastructure automation starts with great infrastructure data.
• Partner directly with customers and with our Solution Architecture/Customer Success teams on L2/L3 escalations—communicating findings, driving root-cause analysis, and resolving complex packaging, deployment, upgrade, and runtime issues across heterogeneous Kubernetes environments. • Drive issues to resolution by reproducing problems locally, isolating root causes, and coordinating fixes with engineering—then documenting learnings in crisp RCAs that become actionable improvements • Build and maintain diagnostics tooling including support bundles, health checks, environment validators, and "what changed?" helpers that make future troubleshooting 10x faster • Own the test automation infrastructure roadmap, improving CI stability, reducing flaky tests, and creating reproducible integration/e2e environments that catch issues before customers do • Establish and maintain performance baselines and regression tests that serve as actionable gates, helping teams catch scale and latency issues early • Improve installation and upgrade robustness by identifying recurring failure modes and eliminating them through product changes, automation, and guardrails • Write production-quality code in Python, Go, or Rust for internal tooling and product improvements that directly enhance reliability • Close the reliability feedback loop by systematically turning field issues into better tests, observability, documentation, and product defaults—measuring success through reduced time-to-resolution and fewer repeat incidents
Product Reliability Engineer
OpsMillGreat infrastructure automation starts with great infrastructure data.
• Partner directly with customers and with our Solution Architecture/Customer Success teams on L2/L3 escalations—communicating findings, driving root-cause analysis, and resolving complex packaging, deployment, upgrade, and runtime issues across heterogeneous Kubernetes environments. • Drive issues to resolution by reproducing problems locally, isolating root causes, and coordinating fixes with engineering—then documenting learnings in crisp RCAs that become actionable improvements • Build and maintain diagnostics tooling including support bundles, health checks, environment validators, and "what changed?" helpers that make future troubleshooting 10x faster • Own the test automation infrastructure roadmap, improving CI stability, reducing flaky tests, and creating reproducible integration/e2e environments that catch issues before customers do • Establish and maintain performance baselines and regression tests that serve as actionable gates, helping teams catch scale and latency issues early • Improve installation and upgrade robustness by identifying recurring failure modes and eliminating them through product changes, automation, and guardrails • Write production-quality code in Python, Go, or Rust for internal tooling and product improvements that directly enhance reliability • Close the reliability feedback loop by systematically turning field issues into better tests, observability, documentation, and product defaults—measuring success through reduced time-to-resolution and fewer repeat incidents
• Collaborate closely with our engineering team to help them be more productive by improving development tools. • Maintain a strong connection to the product and remain comfortable engaging in direct feature development • You will take ownership of CI/CD stack, keeping it robust, observable, stable, and performant. • Research, implement, and develop tools to help developers write new code. • Ensure the CI/CD infrastructure and process is reliable and consistent.
• Assess service maturity and provide insights to development teams • Partner with development teams to implement observability best practices • Enable development teams to become autonomous with their service deployment, support, and infrastructure • Mentor developers on reliability practices, focusing on making them self-sufficient • Act as the bridge, ear and eyes of the Platform Division teams to drive tooling and practice adoption across development teams



