Scalr logo
Scalr

Scalr is a remote state & operations backend for Terraform and OpenTofu.

SRE Tech Lead

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 51-200H1B No SponsorCompany SiteLinkedIn

Location

Ukraine

Posted

5 days ago

Salary

0

Seniority

Senior

Job Description

SRE Tech Lead

Scalr

• Own the Scalr platform's reliability, scalability, and observability • Proactively identify and eliminate risks before they become incidents • Design new architecture components • Promote and enforce SRE and DevOps best practices • Drive strategic technical improvements • Create and maintain the SRE roadmap and backlog • Own the observability strategy end-to-end (monitoring, alerting, dashboards) • Architect and evolve the platform infrastructure • Define and own customer-centric SLIs, SLOs and error budgets • Manage the infrastructure technology stack • Cooperate with other Tech Leads and coordinate interaction with other departments • Lead the resolution of technical challenges

Job Requirements

  • Python (experience in Python scripting is enough)
  • Terraform/OpenTofu
  • Strong knowledge of Linux (RHEL/Debian, bash scripting)
  • Docker
  • Kubernetes
  • Google Cloud Platform
  • Leading SRE teams or initiatives
  • Experience with monitoring and logging tools such as Grafana, Prometheus, Datadog, New Relic, etc.
  • Experience with CI platforms such as GitHub Actions, Drone, CircleCI, etc.
  • Strong written and verbal communication skills
  • Would be a plus: Experience with GitOps, Argo CD, Flux CD or similar
  • Chef, Omnibus, Ruby
  • JavaScript for GitHub Actions

Benefits

  • Attractive compensation and benefits package
  • Long-term contract and tax compensations
  • Flexible schedule and possibility to work entirely remotely
  • Medical insurance
  • 20 working days of paid vacation and 2 weeks of paid sick leaves

Related Categories

Related Job Pages

More DevOps Engineer Jobs

OWKIN logo

Security Engineer – DevSecOps, Code Security

OWKIN

We create the first closed-loop AI generative biology company to create the new standard of biology reasoning

DevOps Engineer5 days ago
Full TimeRemoteTeam 201-500Since 2016

• Conduct in-depth application security assessments and secure code reviews across frontend and backend systems • Partner with engineering teams to remediate vulnerabilities and improve secure coding standards • Review and secure Git-based development workflows and branching strategies • Integrate security controls into CI/CD pipelines in GitHub and DevSecOps processes • Support cloud-native security initiatives across Kubernetes and AWS environments • Use modern application security tooling, including Wiz Code, to identify and prioritise risks • Develop automation and tooling using Python to support security operations and engineering workflows • Advise developers on secure architecture, threat modelling, and security best practices • Collaborate with DevOps, Platform Engineering, and Software Engineering teams to improve overall security posture • Assist with vulnerability management, risk assessment, and remediation tracking • Contribute to security standards, policies, and developer enablement initiatives • On-call rotation for Wiz alerts (paid at an additional rate)

France
Full TimeRemoteTeam 51-200Since 2018

• Co-owner of the architecture: Help drive the architecture and evolution of our cloud infrastructure on Azure and our Kubernetes clusters—designed for high throughput and maximum availability—to support Flip’s rapid global growth. • Drive the resilience strategy: Define our approach to global scaling, zero-downtime deployments, rollback mechanisms, and disaster recovery, ensuring the platform remains available around the clock. • Evolve our observability stack: Optimize our LGTM stack (Loki, Grafana, Tempo, Mimir) into a foundation our engineers can rely on. • Improve our IaC platform: Remove toil at the source and make our infrastructure a true self-service for engineering teams. • Lead during incidents: Take a leading role in major platform incidents, conduct factual post-incident analyses (blameless post-mortems), and turn findings into lasting improvements. • Mentor within the squad: Coach team members, lead RFCs and design reviews, and help engineers grow into stronger SREs. • Shape our roadmap: Collaborate closely with your squad to define the direction of the platform.

Europe
Full TimeRemoteTeam 51-200Since 2018

• Co-own the architecture: Help drive the architecture and evolution of our cloud infrastructure on Azure and our Kubernetes clusters - designed for high throughput and highest availability - to support Flip's rapid growth across the globe. • Drive the resilience strategy: Define how we approach global scaling, zero-downtime deployments, rollback mechanisms and disaster recovery, and make sure the platform stays available around the clock. • Evolve our observability stack: Improve our LGTM stack (Loki, Grafana, Tempo, Mimir) into a foundation our engineers can trust. • Improve our IaC Platform: Eliminate toil at the source, and make our infrastructure truly self-service for engineering teams. • Lead in incidents: Take a leading role in platform-related major incidents, drive blameless post-mortems for the squad, and translate findings into systemic improvements. • Mentor within the squad: Coach teammates, run RFCs and design reviews inside the team, and help engineers grow into stronger SREs. • Shape our roadmap: Partner with your squad to define the platform's direction.

Germany
OpsMill logo

Product Reliability Engineer

OpsMill

Great infrastructure automation starts with great infrastructure data.

DevOps Engineer5 days ago
Full TimeRemoteTeam 11-50Since 2023

• Partner directly with customers and with our Solution Architecture/Customer Success teams on L2/L3 escalations—communicating findings, driving root-cause analysis, and resolving complex packaging, deployment, upgrade, and runtime issues across heterogeneous Kubernetes environments. • Drive issues to resolution by reproducing problems locally, isolating root causes, and coordinating fixes with engineering—then documenting learnings in crisp RCAs that become actionable improvements • Build and maintain diagnostics tooling including support bundles, health checks, environment validators, and "what changed?" helpers that make future troubleshooting 10x faster • Own the test automation infrastructure roadmap, improving CI stability, reducing flaky tests, and creating reproducible integration/e2e environments that catch issues before customers do • Establish and maintain performance baselines and regression tests that serve as actionable gates, helping teams catch scale and latency issues early • Improve installation and upgrade robustness by identifying recurring failure modes and eliminating them through product changes, automation, and guardrails • Write production-quality code in Python, Go, or Rust for internal tooling and product improvements that directly enhance reliability • Close the reliability feedback loop by systematically turning field issues into better tests, observability, documentation, and product defaults—measuring success through reduced time-to-resolution and fewer repeat incidents

United States