Scalr is a remote state & operations backend for Terraform and OpenTofu.
SRE Tech Lead
Location
Ukraine
Posted
5 days ago
Salary
0
Seniority
Senior
Job Description
SRE Tech Lead
Scalr
• Own the Scalr platform's reliability, scalability, and observability • Proactively identify and eliminate risks before they become incidents • Design new architecture components • Promote and enforce SRE and DevOps best practices • Drive strategic technical improvements • Create and maintain the SRE roadmap and backlog • Own the observability strategy end-to-end (monitoring, alerting, dashboards) • Architect and evolve the platform infrastructure • Define and own customer-centric SLIs, SLOs and error budgets • Manage the infrastructure technology stack • Cooperate with other Tech Leads and coordinate interaction with other departments • Lead the resolution of technical challenges
Job Requirements
- Python (experience in Python scripting is enough)
- Terraform/OpenTofu
- Strong knowledge of Linux (RHEL/Debian, bash scripting)
- Docker
- Kubernetes
- Google Cloud Platform
- Leading SRE teams or initiatives
- Experience with monitoring and logging tools such as Grafana, Prometheus, Datadog, New Relic, etc.
- Experience with CI platforms such as GitHub Actions, Drone, CircleCI, etc.
- Strong written and verbal communication skills
- Would be a plus: Experience with GitOps, Argo CD, Flux CD or similar
- Chef, Omnibus, Ruby
- JavaScript for GitHub Actions
Benefits
- Attractive compensation and benefits package
- Long-term contract and tax compensations
- Flexible schedule and possibility to work entirely remotely
- Medical insurance
- 20 working days of paid vacation and 2 weeks of paid sick leaves
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Security Engineer – DevSecOps, Code Security
OWKINWe create the first closed-loop AI generative biology company to create the new standard of biology reasoning
• Conduct in-depth application security assessments and secure code reviews across frontend and backend systems • Partner with engineering teams to remediate vulnerabilities and improve secure coding standards • Review and secure Git-based development workflows and branching strategies • Integrate security controls into CI/CD pipelines in GitHub and DevSecOps processes • Support cloud-native security initiatives across Kubernetes and AWS environments • Use modern application security tooling, including Wiz Code, to identify and prioritise risks • Develop automation and tooling using Python to support security operations and engineering workflows • Advise developers on secure architecture, threat modelling, and security best practices • Collaborate with DevOps, Platform Engineering, and Software Engineering teams to improve overall security posture • Assist with vulnerability management, risk assessment, and remediation tracking • Contribute to security standards, policies, and developer enablement initiatives • On-call rotation for Wiz alerts (paid at an additional rate)
• Co-owner of the architecture: Help drive the architecture and evolution of our cloud infrastructure on Azure and our Kubernetes clusters—designed for high throughput and maximum availability—to support Flip’s rapid global growth. • Drive the resilience strategy: Define our approach to global scaling, zero-downtime deployments, rollback mechanisms, and disaster recovery, ensuring the platform remains available around the clock. • Evolve our observability stack: Optimize our LGTM stack (Loki, Grafana, Tempo, Mimir) into a foundation our engineers can rely on. • Improve our IaC platform: Remove toil at the source and make our infrastructure a true self-service for engineering teams. • Lead during incidents: Take a leading role in major platform incidents, conduct factual post-incident analyses (blameless post-mortems), and turn findings into lasting improvements. • Mentor within the squad: Coach team members, lead RFCs and design reviews, and help engineers grow into stronger SREs. • Shape our roadmap: Collaborate closely with your squad to define the direction of the platform.
• Co-own the architecture: Help drive the architecture and evolution of our cloud infrastructure on Azure and our Kubernetes clusters - designed for high throughput and highest availability - to support Flip's rapid growth across the globe. • Drive the resilience strategy: Define how we approach global scaling, zero-downtime deployments, rollback mechanisms and disaster recovery, and make sure the platform stays available around the clock. • Evolve our observability stack: Improve our LGTM stack (Loki, Grafana, Tempo, Mimir) into a foundation our engineers can trust. • Improve our IaC Platform: Eliminate toil at the source, and make our infrastructure truly self-service for engineering teams. • Lead in incidents: Take a leading role in platform-related major incidents, drive blameless post-mortems for the squad, and translate findings into systemic improvements. • Mentor within the squad: Coach teammates, run RFCs and design reviews inside the team, and help engineers grow into stronger SREs. • Shape our roadmap: Partner with your squad to define the platform's direction.
Product Reliability Engineer
OpsMillGreat infrastructure automation starts with great infrastructure data.
• Partner directly with customers and with our Solution Architecture/Customer Success teams on L2/L3 escalations—communicating findings, driving root-cause analysis, and resolving complex packaging, deployment, upgrade, and runtime issues across heterogeneous Kubernetes environments. • Drive issues to resolution by reproducing problems locally, isolating root causes, and coordinating fixes with engineering—then documenting learnings in crisp RCAs that become actionable improvements • Build and maintain diagnostics tooling including support bundles, health checks, environment validators, and "what changed?" helpers that make future troubleshooting 10x faster • Own the test automation infrastructure roadmap, improving CI stability, reducing flaky tests, and creating reproducible integration/e2e environments that catch issues before customers do • Establish and maintain performance baselines and regression tests that serve as actionable gates, helping teams catch scale and latency issues early • Improve installation and upgrade robustness by identifying recurring failure modes and eliminating them through product changes, automation, and guardrails • Write production-quality code in Python, Go, or Rust for internal tooling and product improvements that directly enhance reliability • Close the reliability feedback loop by systematically turning field issues into better tests, observability, documentation, and product defaults—measuring success through reduced time-to-resolution and fewer repeat incidents



