Job Closed
This listing is no longer active.
Principal DevOps Engineer
Location
Portugal
Posted
127 days ago
Salary
0
Seniority
Lead
Job Description
Principal DevOps Engineer
Intermedia Cloud Communications
• Act as technical lead for DevOps/Platform/Release engineering: set direction, standards, and best practices • Architect and govern end-to-end delivery: infrastructure provisioning, configuration management, CI/CD, release processes, and operations • Design and support Windows-based high availability solutions, with deep ownership of Windows clustering (failover/HA patterns, maintenance, upgrades, troubleshooting) • Lead Linux automation and platform standardization (configuration, patching, hardening, performance tuning) • Own Infrastructure as Code strategy with Terraform (modules, environments, state, governance) • Own automation strategy with Ansible (reusable roles, inventories, secure secrets handling, idempotency) • Build and standardize deployments using Octopus Deploy, GitHub, and Ansible (templates, shared steps, release promotion, rollback) • Design and mature CI/CD pipelines (artifact versioning, approvals, promotion strategy, policy-as-code where applicable) • Establish observability standards using VictoriaMetrics/Prometheus (metrics strategy, alerting, SLO/SLA monitoring, dashboards) • Provide production leadership: incident response, RCA/postmortems, reliability improvements, capacity planning • Mentor engineers, review designs/code, and raise overall engineering quality across teams • Produce and maintain architecture docs, runbooks, and platform roadmaps
Job Requirements
- Bachelors degree in Computer Science or related field
- 7+ years (or equivalent) in DevOps / SRE / Infrastructure Engineering, including leadership in complex environments
- Expert-level experience designing and operating Windows Server HA and clustering (Failover Clustering and related components)
- Strong Linux administration and automation experience (systemd, networking, storage, performance)
- Advanced skills with Terraform and Ansible (architecture, reusable components, secure operations)
- Strong deployment/release engineering experience with Octopus Deploy and GitHub (release governance, environment promotion, rollback)
- Monitoring/observability expertise with VictoriaMetrics and/or Prometheus (alerting strategy, metrics design, operational readiness)
- Production experience running Redis, RabbitMQ, Nginx (HA, tuning, troubleshooting)
- Strong understanding of networking and security fundamentals (TLS, DNS, load balancing, firewalling, least privilege)
- Proven ability to lead cross-team initiatives, make architectural decisions, and communicate clearly
- Kubernetes and container ecosystems (Docker, Helm)
- Nice to have**
- CI/CD platforms beyond GitHub (GitLab CI, Jenkins)
- Logging platforms (ELK/EFK, Loki)
- DR/BCP design, backup automation, zero-downtime upgrade strategies
- PowerShell and advanced scripting, configuration governance, secrets tooling (Vault/SOPS)
- Experience with Virtualization platforms such as VMWare or HyperV
- Experience in building, configuring, and tuning highly available MS SQL Server environments
- Experience in managing VOIP components and protocols (SIP , FreeSwitch, OpenSIP, session border controllers)
- Experience with load balancing components ( F5 LTM, F5 GTM)
- Experience with administering AWS or Azure tenants
Benefits
- We hire, promote, and compensate employees based on their ability to perform their job responsibilities, without regard to race, color, creed, religion, sex, gender, marital status, national origin, ancestry, age, citizenship, physical or mental disability, sexual orientation, or any other basis protected by applicable law (collectively referred to in our Code of Conduct as “Protected Classes”). We do not tolerate employment discrimination in the workplace, and we are committed to making reasonable accommodations for identified disabilities or other limitations as required by all applicable laws. We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.*
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Head of SRE
YousignWe reinvent the eSignature experience with our easy to use, certified and secure solution 🖊️⚡️
• Lead, inspire, and grow a team of 6 SREs with diverse profiles (Infrastructure and Application backgrounds) • Create a unified SRE vision while respecting each profile's specificities and career development needs • Provide coaching and mentorship, including for senior/staff engineers on technical leadership • Manage workload, prioritize and make structured recommendations that will be discussed with Engineering management. • Manage external resources: freelancers and specific expertise as needed • Embody and transmit Yousign's vision, values, and Operating Principles • Drive 90% decommissioning by H1 2026, 100% by EOY 2026 • Define post-migration strategy: resilience (eliminate SPOFs), scalability (support growth), rationalization (observability overhaul) • Drive Platform Engineering vision: IDP/DevHub implementation to reduce cognitive load for product teams • Ensure SLA & SLO for critical B2B trust signature service (thousands of customers, millions of users) • Lead incident management with mature processes (on-call, runbooks, war rooms) • Establish and execute crisis communication plans maintaining trust and transparency • Own Build topics: Automation, CI/CD, technical framing, IDP/DevHub implementation, Developer Experience (self-service, golden paths) • Lead cross-functional initiatives to improve platform efficiency and reliability
Senior Site Reliability Engineer
Latitude.shLatitude.sh is a global bare metal cloud platform built for developers.
• Continuously improve Latitude.sh’s platform reliability and performance • Design, build, and maintain tools to automate operational tasks and incident response • Implement and improve observability solutions, including monitoring, alerting, and tracing • Collaborate with engineering and platform teams to design scalable and resilient systems • Participate in on-call rotations and lead post-incident reviews with a focus on learning • Develop and document processes and runbooks that ensure operational excellence • Contribute to SLOs/SLIs definition and reliability metrics adoption across teams
DevOps Engineer
NaNLABSYour Sidekick for AI, Cloud-Native & Real-Time Data Engineering | Scalable Innovations in Auto, EV, SaaS & Cybersecurity
• Design, implement, and maintain scalable and secure cloud infrastructure in AWS • Manage infrastructure as code using Terraform to ensure consistency, automation, and reliability across environments • Build and improve CI/CD pipelines using GitHub Actions to enable efficient and reliable software delivery • Manage Kubernetes infrastructure (EKS) and support networking, security, and access management within AWS environments • Implement observability practices across services using modern monitoring and alerting tools • Support and scale infrastructure for data and machine learning workloads • Collaborate closely with engineering and data teams to ensure infrastructure supports product scalability and performance • Define and promote DevOps best practices related to automation, reliability, and infrastructure standards • Participate in incident response and help improve reliability through proactive monitoring and infrastructure improvements • Communicate technical decisions, risks, and trade-offs clearly with both technical and non-technical stakeholders
Senior DevOps Engineer
NearsureRemove the barriers to growth by scaling your team fast with top-notch Latin American IT talent
• Design, build, and standardize CI/CD pipelines (Azure DevOps preferred) to support multiple development teams. • Build and evolve a self-service, GitOps-driven platform on Azure. • Contribute to the architecture, design, and implementation of Azure Kubernetes Service (AKS) environments. • Contribute to architecture and security configuration decisions within Azure and Kubernetes. • Implement Infrastructure as Code using Terraform and Helm to standardize and automate cloud environments. • Establish reusable automation patterns and platform standards across teams. • Integrate and automate security tooling within CI/CD pipelines in alignment with the dedicated security team. • Maintain and optimize existing pipelines and platform components. • Standardize CI/CD and deployment practices across 7–8 parallel teams. • Enable development teams through platform improvements and self-service capabilities. • Troubleshoot and debug production issues across services, infrastructure, and Kubernetes clusters. • Participate in incident response and post-mortem activities when required. • Collaborate closely with cross-functional and distributed teams to ensure consistent platform evolution. • Take ownership of platform components, driving improvements proactively and independently.




