Senior Site Reliability Engineer – SRE, DevOps

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 11-50H1B No SponsorCompany SiteLinkedIn

Location

Vietnam

Posted

4 days ago

Salary

0

Seniority

Senior

Job Description

Senior Site Reliability Engineer – SRE, DevOps

qode.world

• Design and operate secure, scalable, and high-quality infrastructure that supports modern applications and advanced AI workloads • Build and maintain robust automation across CI/CD pipelines, infrastructure provisioning, and operational processes to improve reliability and minimize manual effort • Integrate AI-driven solutions into operational workflows to enhance efficiency, detect anomalies, and accelerate delivery • Apply strong systems engineering practices, including monitoring, incident management, performance optimization, and capacity planning • Establish and uphold DevOps best practices, ensuring reproducibility, testing, documentation, and operational excellence • Communicate technical decisions clearly and collaborate cross-functionally to support predictable delivery and effective problem-solving • Provide mentorship and technical leadership, raising the level of platform engineering, DevOps maturity, and overall engineering quality across the organization

Job Requirements

  • 6+ years of progressive experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Engineering.
  • Strong, hands-on experience across multi-cloud environments (AWS, GCP, Azure), including expertise in networking, compute, storage, security, and cost optimization.
  • Deep expertise in containerization and orchestration and extensive experience with Infrastructure as Code (IaC) (e.g., Terraform, Pulumi, CloudFormation).
  • Experience supporting or deploying AI/ML workloads (e.g., model inference, vector databases, GPU workloads), or strong familiarity with the infrastructure requirements for these systems.
  • Proven ability to design, build, and operate highly reliable, scalable production systems utilizing advanced Zero-Downtime Deployment Patterns (e.g., Blue/Green, Canary, progressive delivery, Preview Environments).
  • Expertise in modernizing deployments via GitOps practices (e.g., ArgoCD, Flux) and building Self-Service Developer Platforms that enable engineering efficiency (e.g., environment automation, internal tooling).
  • Experience implementing and managing Multi-Cloud API Gateways and Edge Routing solutions.
  • Strong background in platform security, including secrets management, Identity and Access Control (IAM), and Runtime/Security Hardening.
  • Solid understanding and practical experience with modern observability stacks.
  • Excellent communication and collaboration skills with a proven ability to describe complex infrastructure decisions clearly and a background in mentoring engineers and driving improvements in engineering practices.
  • Familiarity with modern programming languages like Node.js, NestJS, and Python is highly desirable for extending DevOps capabilities or integrating tooling.

Benefits

  • Attractive salary range and open to negotiate for strong fits.
  • Hybrid/Remote-friendly culture. Work where you grow best!
  • Flexible hours, async teamwork. Focus time is respected.
  • Work equipment support.
  • Allowance for certification & skill development.
  • Year-end bonus & performance-based rewards.
  • 22 paid leaves from your 5th year. Take a full month off.
  • Career growth with personal coaching sessions.
  • Open, collaborative team culture. No micromanagement, only trust.
  • Tools & AI-powered workflows that make remote work easier.

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Empower AI logo

Senior DevOps Engineer

Empower AI

Empower AI (formerly NCI) elevates public sector teams with the power of AI, to ensure America’s missions are met.

DevOps Engineer4 days ago
Full TimeRemoteTeam 501-1,000Since 1989H1B No Sponsor

• Mentor fellow engineers • Architect scalable infrastructure solutions • Build automated solutions • Respond to infrastructure support requests • Triaging and investigating complexities • Execute OS hardening and deploying infrastructure with Terraform • Communicate effectively with peers • Participate in agile ceremonies • Support patching operations on weekends

United States
Full TimeRemoteTeam 201-500Since 2017H1B No Sponsor

• design and build the infrastructure primitives that define how our CI/CD platform, build systems, and developer environments scale across the entire engineering org. • help build and operate the Kubernetes-based control plane behind our CI/CD platform, including: - GitHub Actions self-hosted runner infrastructure (autoscaling, isolation, cost/perf tuning) - GitHub Apps and GitHub-as-code (permissions, webhooks, org-wide automation) - Secure network access for CI/CD and remote dev environments (Tailscale) - GitOps-driven deployment of platform services (Flux) - Ephemeral/on-demand developer environments and build systems • develop the core infrastructure components — including Kubernetes Operators and scaling automation — that product teams adopt directly, reducing bespoke per-team CI/CD and environment tooling. • building the systems that define how engineering teams build, test, and deploy, shaping the reliability and scalability of the developer experience org-wide.

Canada
$90K - $125K / year
AITASTIC AG logo

DevOps Engineer

AITASTIC AG

Next Level Consumer and Communication Insights

DevOps Engineer4 days ago
Full TimeRemoteTeam 51-200H1B No Sponsor

• Operate and deploy applications within an SOA architecture (GCP, Kubernetes, Terraform) • Ensure high availability of databases (Elasticsearch, Redis, and vector databases) • Optimize logging and monitoring using Grafana and Prometheus • Contribute to the design and implementation of security concepts

Germany

Role Description We’re looking for a DevOps Engineer to join our fully-remote team and help architect, deploy, and operate mission-critical cloud platforms that support everything from financial services to healthcare and government systems. This is a hands-on role where you’ll be working at the intersection of automation, Kubernetes, GitOps, and sovereign cloud infrastructure — helping organisations transition from legacy systems to resilient, containerized environments built for performance, security, and data sovereignty. If you're passionate about CI/CD, infrastructure-as-code, and building scalable deployment workflows that actually make developers’ lives easier — you’ll feel right at home here. What You’ll Be Working On - Designing and maintaining automated deployment pipelines and cloud-native platforms running on fully sovereign infrastructure. - Supporting clients through: - Infrastructure migrations - Platform modernisation - High-availability cluster deployments - Secure containerised workloads - Observability and incident response - Working closely with development teams to streamline delivery pipelines and implement scalable deployment strategies across multi-tenant OpenShift platforms. Your Impact - Design, implement, and maintain CI/CD pipelines to streamline build and release processes. - Install, configure, and manage OpenShift clusters across staging and production environments. - Deploy and manage storage platforms including: - OpenShift Data Foundation (ODF) - Ceph - Optimise Kubernetes, Docker, and OpenShift environments for performance, availability, and security. - Automate infrastructure provisioning using: - Terraform - Ansible - FluxCD - Develop Helm charts and Kustomize configurations within a GitOps-driven deployment model. - Monitor system performance and troubleshoot production incidents. - Implement monitoring and logging solutions using: - Prometheus - ELK Stack - Deploy and manage security integrations: - TLS Certificates - Authentication via Keycloak - ISO-compliant security frameworks - Manage containerised database environments: - PostgreSQL - MySQL - Oversee network, storage, and data centre integrations. - Automate recurring infrastructure tasks and optimise platform workflows. - Support customers directly across planning, deployment, and operational phases. - Document systems and contribute to internal knowledge-sharing initiatives. Qualifications - Solid experience with modern CI/CD practices. - Hands-on expertise in: - Kubernetes - Docker - Helm - Kustomize - OpenShift - Experience with GitOps platforms: - FluxCD - ArgoCD - Familiarity with storage platforms such as: - OpenShift Data Foundation - Ceph - Infrastructure-as-Code experience with: - Terraform - Ansible - Scripting in: - Bash - Go or Python - Experience managing relational databases: - PostgreSQL - MySQL - Experience working with cloud and on-premise environments. - Monitoring and logging with: - Prometheus - ELK Stack - Strong familiarity with: - Git - Jira - Confluence - Knowledge of networking, storage, and data centre operations. - Experience with security standards, TLS, and ISO compliance. - Excellent communication skills and the ability to work independently in a remote environment. Benefits - Work on enterprise-scale infrastructure running on cutting-edge hardware — including NVIDIA GPU clusters supporting AI training, inference, and HPC workloads. - Ensure every deployment meets the highest standards of Swiss data protection and sovereignty. - Empowered to propose, build, and ship solutions that balance innovation with real-world compliance and operational resilience.

Worldwide