Wand AI logo
Wand AI

Your Future Workforce

Staff Site Reliability Engineer

DevOps EngineerDevOps EngineerFull TimeRemoteLeadTeam 51-200Since 2022Company SiteLinkedIn

Location

Europe

Posted

10 hours ago

Salary

0

Seniority

Lead

Job Description

Staff Site Reliability Engineer

Wand AI

• Architect, deploy, and operate scalable, secure production environments (AWS preferred). • Lead reliability improvements across multiple engineering streams. • Design and evolve Kubernetes-based infrastructure, including migration and optimisation initiatives. • Build and enforce strong Infrastructure-as-Code standards. • Define and operationalise SLIs, SLOs, and error budgets. • Strengthen observability across applications, infrastructure, data pipelines, and ML systems. • Work closely with product and data teams to integrate model analytics and product telemetry into reliability insights. • Work across and optimise the entire CI/CD pipeline, from build to deploy to rollback. • Improve release safety, deployment frequency, and predictability of SLAs. • Lead incident response for complex cross-system failures and drive postmortems. • Reduce operational toil through automation and platform engineering improvements. • Design processes and tooling to absorb, standardise, and troubleshoot customer environments. • Support and productionise ML workloads (MLOps practices including model deployment, monitoring, retraining workflows). • Ensure infrastructure aligns with enterprise-grade security and regulatory requirements. • Mentor engineers and raise the overall reliability bar across teams.

Job Requirements

  • Extensive hands-on experience in SRE or Production Engineering roles.
  • Demonstrated experience building or scaling SRE practices in high-growth or complex environments.
  • Deep expertise in AWS or Azure-based cloud infrastructure.
  • Strong experience with Kubernetes (including migration, scaling, and production hardening).
  • Advanced Infrastructure-as-Code experience (Terraform or equivalent).
  • End-to-end CI/CD pipeline design and optimisation experience.
  • Strong experience with observability tooling across distributed systems.
  • Experience troubleshooting complex multi-tenant or customer-hosted environments.
  • Experience supporting production data platforms and ML systems.
  • MLOps experience, including model deployment and monitoring.
  • Strong understanding of distributed systems, scalability, and fault tolerance.
  • Systems thinker who understands interactions across infrastructure, product, data, and ML.
  • Excellent communication skills and ability to work cross-functionally.

Benefits

  • Health insurance
  • Paid time off
  • Professional development

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Wand AI logo

Head of SRE

Wand AI

Your Future Workforce

DevOps Engineer10 hours ago
Full TimeRemoteTeam 51-200Since 2022

• Own and lead all SRE-related strategy, standards, and execution. • Embed SRE culture and operational excellence across engineering teams. • Review the current infrastructure and operational model; redesign and rebuild where needed. • Architect, deploy, and maintain scalable, secure production environments. • Define and implement SLIs, SLOs, and uptime targets. • Establish robust monitoring, alerting, and observability practices. • Design and implement incident management, RCA and postmortem processes. • Build and manage sustainable on-call frameworks and escalation models. • Automate the software delivery lifecycle to improve release predictability and safety. • Create reproducible environments and IaaC provisioning templates. • Improve system performance, availability, and reliability. • Support and productionise data platforms and ML workloads. • Partner closely with QA and Engineering leadership to improve release quality and stability. • Ensure infrastructure meets enterprise-grade security and regulatory requirements. • Hire, manage, and mentor a team of SRE engineers.

Europe
Scalr logo

SRE Tech Lead

Scalr

Scalr is a remote state & operations backend for Terraform and OpenTofu.

DevOps Engineer10 hours ago
Full TimeRemoteTeam 51-200H1B No Sponsor

• Own the Scalr platform's reliability, scalability, and observability • Proactively identify and eliminate risks before they become incidents • Design new architecture components • Promote and enforce SRE and DevOps best practices • Drive strategic technical improvements • Create and maintain the SRE roadmap and backlog • Own the observability strategy end-to-end (monitoring, alerting, dashboards) • Architect and evolve the platform infrastructure • Define and own customer-centric SLIs, SLOs and error budgets • Manage the infrastructure technology stack • Cooperate with other Tech Leads and coordinate interaction with other departments • Lead the resolution of technical challenges

Ukraine
OWKIN logo

Security Engineer – DevSecOps, Code Security

OWKIN

We create the first closed-loop AI generative biology company to create the new standard of biology reasoning

DevOps Engineer10 hours ago
Full TimeRemoteTeam 201-500Since 2016

• Conduct in-depth application security assessments and secure code reviews across frontend and backend systems • Partner with engineering teams to remediate vulnerabilities and improve secure coding standards • Review and secure Git-based development workflows and branching strategies • Integrate security controls into CI/CD pipelines in GitHub and DevSecOps processes • Support cloud-native security initiatives across Kubernetes and AWS environments • Use modern application security tooling, including Wiz Code, to identify and prioritise risks • Develop automation and tooling using Python to support security operations and engineering workflows • Advise developers on secure architecture, threat modelling, and security best practices • Collaborate with DevOps, Platform Engineering, and Software Engineering teams to improve overall security posture • Assist with vulnerability management, risk assessment, and remediation tracking • Contribute to security standards, policies, and developer enablement initiatives • On-call rotation for Wiz alerts (paid at an additional rate)

France
DevOps Engineer10 hours ago
Full TimeRemoteTeam 51-200Since 2018

• Co-owner of the architecture: Help drive the architecture and evolution of our cloud infrastructure on Azure and our Kubernetes clusters—designed for high throughput and maximum availability—to support Flip’s rapid global growth. • Drive the resilience strategy: Define our approach to global scaling, zero-downtime deployments, rollback mechanisms, and disaster recovery, ensuring the platform remains available around the clock. • Evolve our observability stack: Optimize our LGTM stack (Loki, Grafana, Tempo, Mimir) into a foundation our engineers can rely on. • Improve our IaC platform: Remove toil at the source and make our infrastructure a true self-service for engineering teams. • Lead during incidents: Take a leading role in major platform incidents, conduct factual post-incident analyses (blameless post-mortems), and turn findings into lasting improvements. • Mentor within the squad: Coach team members, lead RFCs and design reviews, and help engineers grow into stronger SREs. • Shape our roadmap: Collaborate closely with your squad to define the direction of the platform.

Europe