IT Resources is a specialized staffing firm focused on connecting skilled technology professionals with companies that need high-impact talent. Our recruiters work on real IT opportunities, partner closely with experienced account managers, and have the freedom and support to build lasting candidate relationships. Location Requirement This position is fully remote, but applicants must currently reside in Florida. Candidates located outside Florida will not be considered.
Senior Site Reliability Engineer
Location
Vietnam
Posted
9 days ago
Salary
0
Seniority
Senior
No structured requirement data.
Job Description
Senior Site Reliability Engineer
Qode
Role Description As a Senior DevOps Engineer, you will work closely with Product, Engineering, and AI teams to shape our infrastructure strategy, design resilient cloud architectures, and ensure our platforms are secure, scalable, and high-performing. You will play a key role in bringing AI systems into production, enabling reliable delivery, strong observability, and operational excellence across our products and internal systems. - Design and operate secure, scalable, and high-quality infrastructure that supports modern applications and advanced AI workloads. - Build and maintain robust automation across CI/CD pipelines, infrastructure provisioning, and operational processes to improve reliability and minimize manual effort. - Integrate AI-driven solutions into operational workflows to enhance efficiency, detect anomalies, and accelerate delivery. - Apply strong systems engineering practices, including monitoring, incident management, performance optimization, and capacity planning. - Establish and uphold DevOps best practices, ensuring reproducibility, testing, documentation, and operational excellence. - Communicate technical decisions clearly and collaborate cross-functionally to support predictable delivery and effective problem-solving. - Provide mentorship and technical leadership, raising the level of platform engineering, DevOps maturity, and overall engineering quality across the organization. Qualifications - 6+ years of progressive experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Engineering. - Strong, hands-on experience across multi-cloud environments (AWS, GCP, Azure), including expertise in networking, compute, storage, security, and cost optimization. - Deep expertise in containerization and orchestration and extensive experience with Infrastructure as Code (IaC) (e.g., Terraform, Pulumi, CloudFormation). - Experience supporting or deploying AI/ML workloads (e.g., model inference, vector databases, GPU workloads), or strong familiarity with the infrastructure requirements for these systems. - Proven ability to design, build, and operate highly reliable, scalable production systems utilizing advanced Zero-Downtime Deployment Patterns (e.g., Blue/Green, Canary, progressive delivery, Preview Environments). - Expertise in modernizing deployments via GitOps practices (e.g., ArgoCD, Flux) and building Self-Service Developer Platforms that enable engineering efficiency (e.g., environment automation, internal tooling). - Experience implementing and managing Multi-Cloud API Gateways and Edge Routing solutions. - Strong background in platform security, including secrets management, Identity and Access Control (IAM), and Runtime/Security Hardening. - Solid understanding and practical experience with modern observability stacks. - Excellent communication and collaboration skills with a proven ability to describe complex infrastructure decisions clearly and a background in mentoring engineers and driving improvements in engineering practices. - Familiarity with modern programming languages like Node.js, NestJS, and Python is highly desirable for extending DevOps capabilities or integrating tooling.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Role Description Gestalte moderne Infrastruktur mit uns. Werde Teil der Contensi Software GmbH. Zur Verstärkung unseres Teams suchen wir eine:n Senior DevOps Engineer, der/die mit fundierter Erfahrung, Weitblick und Leidenschaft für Technologie den Unterschied macht. - Du arbeitest eng mit unseren Kunden aus Industrie, öffentlichem Sektor und Wissenschaft an anspruchsvollen Infrastrukturprojekten, von der Planung bis zum stabilen Betrieb. - Du übernimmst die Architektur, Einrichtung und Weiterentwicklung von Cloud- und OnPrem-Infrastrukturen, darunter Public-Cloud-Umgebungen wie AWS, Azure, GCP sowie Plattformen wie Kubernetes und OpenShift. - Du betreibst und wartest bestehende Systeme und Rechenzentrumsinfrastrukturen, inklusive Hybrid- und Multi-Cloud-Setups. - Du analysierst bestehende Architekturen und unterstützt bei der Migration und Modernisierung von Legacy-Umgebungen. - Du richtest GitOps- und CI/CD-Workflows ein (z. B. mit ArgoCD, GitLab CI, Azure DevOps) und begleitest unsere Kunden bei der Automatisierung ihrer Entwicklungs- und Betriebsprozesse. - Du berätst zu aktuellen Technologien und Best Practices, insbesondere im Bereich Containerisierung, Cloud-native Tools und Infrastructure-as-Code (Terraform, Ansible). - Du arbeitest konzeptionell mit an Themen wie Netzwerkdesign, Storage-Architekturen (z. B. Ceph, ODF, NFS) und Security-Standards (BSI, ISO27001, IAM, PKI). - Bei Wunsch betreust du Kunden auch vor Ort in Deutschland – und unterstützt unsere Kolleg:innen durch dein Fachwissen im Team. Qualifications - Mindestens 6 Jahre Berufserfahrung im Bereich DevOps, Infrastruktur oder Site Reliability Engineering. - Erfahrung mit Cloud-Plattformen: AWS, Azure, GCP. - Sehr gute Kenntnisse in Kubernetes und OpenShift. - Praktische Erfahrung mit einem oder mehreren CI/CD-Tools: Jenkins, GitLab CI, GitHub Actions, Azure DevOps o. ä. - Kenntnisse in Logging- und Monitoring-Systemen wie ELK, Datadog, Prometheus, New Relic o. ä. - Sehr gute Kenntnisse in Infrastructure as Code (Terraform, Ansible, Puppet). - Solides Netzwerkverständnis (TCP/IP, DNS, Routing, Firewalls etc.). - Erfahrung mit Storage-Systemen wie Ceph, MinIO etc. - Know-how in Security-Standards & Compliance (z. B. BSI, ISO27001, PKI, IAM). - Fundierte Kenntnisse in mindestens einer Programmiersprache: z. B. Java, Python, Go, Node.js oder TypeScript. - Du bist motiviert, dich kontinuierlich weiterzuentwickeln und neue Technologien zu erlernen. - Fliessende Deutsch- und Englischkenntnisse. Benefits - Spannende Kundenprojekte mit technologischer Tiefe. - Remote-Arbeit mit gelegentlichen vor Ort Besuchen beim Kunden. - Gestaltungsfreiraum - bring deine Ideen ein, übernimm Verantwortung und wachse mit uns. - Weiterbildungsbudget für Zertifikate, Konferenzen und Fachliteratur. - Modernstes Equipment (MacBook Pro & iPhone zur privaten/beruflichen Nutzung). - Zugang zu Lernplattformen wie Pluralsight und O’Reilly Online Learning. - Ergonomisches Homeoffice-Setup nach Wunsch. - Wellpass-Membership Zuschuss für sportliche Aktivitäten wie Schwimmbäder, Fitness-Studios uvm. - 30 Tage Urlaub.
• Administration and support of service operations • Setting up build and code delivery processes (CI/CD) • Automation and optimization of routine tasks
• Manage and maintain Copado platform across multiple product teams • Design and optimize CI/CD pipelines • Automate manual tasks such as promotions and security scans • Provide L2/L3 support and troubleshooting • Onboard new teams and define best practices • Cross-train internal team members and collaborate across teams
DevOps Engineer
SOCOTEC UK & IrelandSOCOTEC’s Environment and Safety team provides expert environmental, health and safety consultancy and compliance services that help organisations manage risk, protect people and safeguard the environment. Our services include: Environmental monitoring and consultancy Water safety and hygiene solutions with Legionella risk assessments and water system management Fire safety consultancy and inspections Occupational hygiene assessments and workplace exposure monitoring Specialist advisory support on regulatory and sustainability challenges We use industry-leading tools and science-based approaches to deliver accurate data, practical guidance and innovative solutions that ensure compliance with UK legislation, minimise environmental impact and create healthier, safer workplaces for clients.
Role Description Are you passionate about infrastructure, containers, and modern platform technologies? This could be your opportunity to build a rewarding career as an Infrastructure Engineer, while playing a vital role in keeping SOCOTEC UK's technology estate running and evolving. As our business continues to grow and expand across the UK, our IT Operations team is growing with it. We have an exciting opportunity for a driven and technically curious DevOps Engineer to join our team and wear the SOCOTEC badge with pride. You will work across a broad and modern technology stack — from Linux and Microsoft through to Kubernetes, Rancher, and our observability platform — giving you genuine exposure and development opportunities that are hard to find in a single role. We are looking for a motivated, hands-on engineer who is ready to roll their sleeves up and get stuck in. You will contribute to the day-to-day running of our infrastructure, support our containerised workloads, and help maintain the monitoring and alerting tools that keep our services healthy. Whether you come with experience across the full stack or strong foundations you are eager to build on, you will be supported by an experienced team who are invested in your growth. You will embody our core behaviours of integrity, curiosity, warmth, and ambition. As part of the IT Operations team, it is key that you can work independently and take ownership of tasks, whilst also collaborating effectively as part of a close-knit team. If you are ready to take the next step in your infrastructure career with a well-established, UK-wide organisation, we would love to hear from you. Qualifications - Demonstrable, hands-on experience with Linux server administration in a production environment - Strong understanding of containerisation — Docker and Kubernetes — and experience managing workloads in a Kubernetes cluster - Hands-on experience with Rancher for Kubernetes cluster management and lifecycle - Experience configuring and maintaining Prometheus and Grafana for monitoring and observability - Experience with VMware vSphere / ESXi for virtualisation management - Good working knowledge of the Microsoft stack: Active Directory, Entra ID / Azure AD, Microsoft 365, and Azure - Ability to write scripts and automate tasks (Bash, PowerShell, or Python) Requirements - Assist in infrastructure and application deployment using Infrastructure as Code (IaC) tools, particularly Terraform - Implement, maintain, and develop the company’s modern application hosting environment, including RKE2, Rancher, Longhorn, Docker, and source control (Git) - Provide third-line support for the Kubernetes (K8s) environment and supporting infrastructure and management tools, including reacting to observability alerts and reports - Implement, maintain, and develop the Observability platform, including Grafana, Prometheus, Loki, and Alert Manager - Provision, configure, and maintain Ubuntu servers; investigate Linux-level performance, service, logging, disk, memory, and connectivity issues - Assist with monitoring and alert management for infrastructure and applications - Provide 2nd/3rd line technical support for infrastructure-related issues Benefits - Competitive salary - 25 days holiday with the opportunity to buy more - An electric car scheme (where applicable) - Employee recognition schemes - Family friendly support - Employee benefits and discounts app - Employee assistance programmes - Enhanced company pension

