Senior DevOps/ Site Reliability Engineer

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 201-500

Location

Vietnam

Posted

50 days ago

Salary

0

Seniority

Senior

No structured requirement data.

Job Description

Senior DevOps/ Site Reliability Engineer

GRADION

Role Description Gradion is expanding its SRE team for a client with a long-term managed services. You will be part of a global, follow-the-sun SRE function, responsible for platform stability, cloud infrastructure, and incident response. This role suits engineers who are technically solid, self-directed, and comfortable operating in a fast-moving, internationally distributed environment. You will go through a structured onboarding alongside an internal SRE team before taking on independent operational responsibility. What You Will Do - Own platform availability: monitor, triage, and resolve incidents within defined SLA windows - Manage cloud infrastructure on AWS and/or GCP - provisioning, scaling, and day-to-day operations - Maintain and improve CI/CD pipelines and GitOps workflows - Operate observability systems: monitoring, logging, and alerting at production scale - Participate in on-call rotation as part of the global follow-the-sun coverage model - Configure, deploy, and manage AI tooling and MCP servers in production environments - Contribute to infrastructure automation, scripting, and internal tooling - Write clear post-incident reviews and contribute to the monthly operational report - Collaborate closely with engineering teams across multiple time zones Qualifications - 4+ years in a DevOps / SRE / Platform Engineering role within an international team - Solid Kubernetes knowledge - cluster operations, troubleshooting, and configuration - Hands-on cloud experience with AWS and/or GCP - Good understanding of networking fundamentals - DNS, load balancing, firewalls, VPC - Scripting and automation skills (Python, Bash, or similar) - Experience with CI/CD tools and GitOps-based delivery - Working knowledge of monitoring and observability systems (Prometheus, ELK, or equivalent) - Good English - daily communication with European stakeholders is a core requirement - Self-directed and proactive - you ask the right questions and drive issues to resolution without waiting to be told Nice to Have - Experience configuring and managing MCP servers and AI tooling in production - Exposure to AI enablement workflows or LLM infrastructure - Background supporting eCommerce or SaaS platforms - Familiarity with the Frontastic / commercetools Frontend ecosystem Benefits - Join Vietnam’s Best IT Company - Gradion Vietnam was recognized by ITViec for 8 consecutive years, including 2 successive years as the Winner. - Career Growth & Leadership Development - Work closely with our leadership team, gain mentorship from experienced executives, and have direct exposure to high-level strategic decisions. - AI-First Engineering & Strategic Consulting - Our engineering culture integrates AI as a core driver of design, development, and optimization. - Competitive Compensation - Expect an attractive salary, performance-based bonuses, and a benefits package that reflects your impact. - Performance bonus of up to 2 months’ salary. - Performance review twice a year, so your growth is recognized and rewarded. - Premium healthcare for you, plus an annual health check. - 15 days of annual leave. - Full salary during probation. - Hybrid working for real flexibility. - Monthly Happy Hour and Community Tech activities. - Work on global projects as part of an innovation team that shapes ideas for the hi-tech world. - Diverse training programs to keep you growing. Working Time Monday - Friday (9 AM - 6 PM) Location - Ho Chi Minh office: Podium Floor, Sapphire 2 tower, 92 Nguyen Huu Canh Street, Thanh My Tay Ward, Ho Chi Minh City, Vietnam. - Da Nang office: 23rd Floor, G8 Golden Building, 65 Hai Phong, Hai Chau Ward, Da Nang City, Vietnam. - Remote Work: Candidates based in Hanoi or Can Tho are welcome to work remotely.

Related Categories

Related Job Pages

More DevOps Engineer Jobs

SimScale logo

Senior SRE / Platform Engineer

SimScale

SimScale empowers engineers with AI-native cloud simulation, enabling thousands of engineering decisions in seconds.

DevOps Engineer50 days ago
Full TimeRemoteTeam 51-200Since 2012

• Evolve our Kubernetes platform: Evaluate and adopt technologies such as Kubernetes Gateway API and service mesh patterns, and coordinate platform evolution across 10+ engineering teams. • Take observability to the next level: Drive organization-wide adoption of OpenTelemetry for distributed tracing and metrics, and help teams define meaningful SLOs. • Shape multi-region architecture and data residency: Support our move from an EU-centered footprint toward a global, multi-cloud architecture that satisfies disaster-recovery and data-residency requirements. • Own cloud cost and efficiency at scale: Keep petabyte-scale infrastructure cost-efficient, secure, and well-instrumented. • Improve tooling: Build self-service AWS account provisioning, guardrails and AI-assisted automations that help engineering teams manage infrastructure safely and efficiently at scale.

Germany
SoftExpert - Software for Excellence logo

Senior DevOps

SoftExpert - Software for Excellence

Software all-in-one para gestão da transformação digital, inovação e conformidade.

DevOps Engineer50 days ago
Full TimeRemoteTeam 501-1,000Since 1995H1B No Sponsor

• Design, implement, and maintain AWS architectures focused on availability, scalability, and security. • Administer and optimize services such as EC2, EKS, RDS, S3, EFS, Lambda, CloudFront, and other platform resources. • Plan and execute migrations, upgrades, and infrastructure changes with minimal operational impact. • Manage complex network environments using VPCs, Transit Gateway, Site-to-Site VPN, Peering, and routing. • Configure and maintain Security Groups, NACLs, and access policies following security best practices. • Investigate and resolve connectivity, performance, and routing issues. • Implement and manage security and protection solutions in cloud environments. • Analyze and respond to security alerts, vulnerabilities, and risks. • Ensure infrastructure compliance through hardening, access management, and periodic reviews. • Administer Linux servers in production environments. • Perform patching, hardening, and lifecycle management of instances. • Create and maintain custom images for different use cases. • Develop and maintain infrastructure using Terraform. • Automate provisioning and configurations with Ansible. • Build and maintain CI/CD pipelines for infrastructure and applications. • Develop scripts to automate operational tasks. • Manage Kubernetes clusters (EKS), including upgrades, troubleshooting, security, and scalability. • Create and optimize Docker images. • Implement security policies and access control within clusters. • Monitor environments using observability tools. • Participate in incident management and production issue resolution. • Conduct root cause analyses (RCA) and implement preventive actions.

Brazil
In All Media logo

DevOps Engineer – Cloud

In All Media

Imagine the future of business. Ideas for a Digital Renaissance.

DevOps Engineer50 days ago
Full TimeRemoteTeam 1,001-5,000H1B No Sponsor

• Lead Cloud Migration: Spearhead the migration of workloads from AWS to Azure using an AKS-first (Azure Kubernetes Service) model. • Infrastructure as Code: Design and implement robust IaC for Azure environments. • Pipeline Standardization: Migrate, optimize, and standardize CI/CD pipelines. • Architecture Modernization: Refactor legacy infrastructure into cloud-native Azure architectures. • Risk Mitigation: Ensure minimal disruption during migration by designing efficient downtime and rollback plans. • Environment Optimization: Continuously optimize cloud environments for cost, performance, and scalability. • Cross-Functional Collaboration: Partner with engineering teams to modernize deployment workflows. • Documentation: Document migration patterns and create reusable templates for the broader team.

Brazil

DevSecOps Engineer SR

Encora Digital

Encora, a leader in digital engineering, drives innovation by crafting cutting-edge, cloud-first, data-first, and AI-first solutions that redefine industries. S

DevOps Engineer51 days ago

Role Description We at Coforge are hiring a DevSecOps Engineer SR with the following skill set. - Design, implement, and maintain secure CI/CD pipelines for frontend, backend, and infrastructure delivery. - Automate infrastructure provisioning, configuration, and deployment processes for cloud and on-prem compatible environments. - Establish security controls across the software delivery lifecycle, including secrets management, dependency scanning, container security, and pipeline guardrails. - Implement secure perimeter and runtime patterns for the external portal, including API gateway controls, network segmentation, identity integration, and access policies. - Support deployment architectures for a scalable platform, secure external authentication, file distribution mechanisms, notifications, audit logging, and portal-to-enterprise platform integrations. - Define and maintain observability capabilities such as logs, metrics, alerting, and health monitoring for platform services and integrations. - Partner with engineering teams to embed security-by-design and operational best practices into application development. - Contribute to incident response readiness, vulnerability remediation workflows, compliance-aligned controls, and environment hardening. - Document deployment standards, platform architecture, runbooks, and operational procedures. Qualifications - Strong experience with DevOps and DevSecOps practices across build, release, deployment, and runtime operations. - Hands-on experience with CI/CD pipeline design, infrastructure as code, automation scripting, and environment provisioning. - Strong knowledge of cloud and hybrid deployment architectures, containerization, and orchestration concepts. - Solid understanding of application and infrastructure security, secrets management, IAM, vulnerability scanning, and secure delivery controls. - Experience implementing observability, alerting, logging, audit, and platform monitoring practices. - Familiarity with network security, gateway patterns, secure external access, and production hardening approaches. - Experience collaborating with software engineering teams to support secure and reliable software delivery. - Strong troubleshooting, documentation, and operational problem-solving skills. Requirements - Experience with Docker, Kubernetes, and secure container runtime practices. - Experience with cloud platforms such as AWS and secure storage/file distribution patterns. - Knowledge of SAST, DAST, dependency scanning, policy-as-code, and security compliance frameworks. - Experience supporting systems with multi-tenant isolation and external security requirements. - Familiarity with Python-based application environments, API-centric architectures, and enterprise integration patterns. Posted On 12-06-2026 At Coforge, we hire professionals based solely on their skills and do not discriminate based on age, disability, religion, gender, sexual orientation, socioeconomic status, or nationality.

Brazil
Job Closed