Job Closed
This listing is no longer active.
We help the best professionals and companies find each other despite borders.
Senior DevOps Engineer – Highload, Cloud, Data-Intensive Systems
Location
Spain
Posted
159 days ago
Salary
€5K - €8K / month
Seniority
Senior
Job Description
Senior DevOps Engineer – Highload, Cloud, Data-Intensive Systems
Alex Staff Agency
• Operation of production services and infrastructure (server provisioning/decommissioning, updates, replacements, performance troubleshooting) • Support and development of Infrastructure as Code (Terraform / Ansible: modules, roles, standards, reviews) • Monitoring, alerting, backups, and regular recovery checks • Development of service and infrastructure automation • Development of CI/CD and release procedures • Incident diagnosis and resolution, support for product teams • Traffic analytics, bot and attack protection tools • Responsibility for 24/7 platform stability
Job Requirements
- 4+ years of experience operating Linux/Ubuntu infrastructure and production services
- Strong understanding of networking and troubleshooting
- Kubernetes (cluster operations), Rancher, Docker / containerd
- Hands-on experience with Ansible and Terraform
- Monitoring: Prometheus / Thanos / Telegraf / Grafana / Sentry
- CI/CD: Jenkins
- Automation: Bash, Python
- Experience working with LVM
- Nice to have: Experience working with blockchain nodes
- Diagnosis and tuning of ClickHouse and MongoDB in high-load clusters
- Providers: Hetzner / OVHcloud
- Cloudflare (edge, DDoS), experience with AWS
- Handling abuse tickets with hosting providers
Benefits
- 5,000 – 8,000 € net
- Format: office / hybrid / remote
- Location: Spain (Barcelona and suburbs) or remote (CET ±2)
- Full-time
- Opportunity to genuinely influence architecture and processes
- Mature engineering team and reasonable expectations
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Write, configure, and deploy code that improves service reliability for existing or new systems; set standard for others with respect to code quality. • Provide helpful and actionable feedback and review for code or production changes. • Drive repair/optimization of complex systems with consideration towards a wide range of contributing factors. • Lead debugging, troubleshooting, and analysis of service architecture and design. • Participate in on-call rotation. • Write documentation: design, system analysis, runbooks, playbooks. Provide design feedback and uplevel design skills of others. • Implement and manage SRE monitoring applications using AI, Python, and Observability data. • Develop tooling using Terraform and other IaC tools to ensure visibility and proactive issue detection across our platforms. • Work within GCP infrastructure, optimizing performance, and cost, and scaling resources to meet demand. • Collaborate with development teams to enhance system reliability and performance, applying a platform engineering mindset to system administration tasks. • Develop and maintain AI-enhanced automated solutions for operational aspects such as on-call monitoring, performance tuning, and disaster recovery. • Troubleshoot and resolve issues in our dev, test, and production environments. • Participate in postmortem analysis and create preventative measures for future incidents. • Implement and maintain security best practices across our infrastructure, ensuring compliance with industry standards and internal policies. Participate in security audits and vulnerability assessments. • Participate in capacity planning and forecasting efforts to ensure our systems can handle future growth and demand. Analyze trends and make recommendations for resource allocation. • Identify and address performance bottlenecks through code profiling, system analysis, and configuration tuning. Implement and monitor performance metrics to proactively identify and resolve issues. • Develop, maintain, and test disaster recovery plans and procedures to ensure business continuity in the event of a major outage or disaster. Participate in regular disaster recovery exercises. • Contribute to internal knowledge bases and documentation.
Senior DevOps Engineer – AWS, Azure
NetguruNetguru builds software that lets people do things differently.
• Architect and operate Kubernetes-based deployments across AWS and Azure • Design, implement, and maintain automated build processes and scripts for Node.js, Python, Java, Flutter, React, and Angular projects. • Develop and manage containerization workflows using Docker, including multi-stage builds, optimized images, and security best practices. • Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, CircleCI, and other relevant tooling. • Implement GitOps workflows using platforms such as ArgoCD or Flux, ensuring declarative infrastructure and automated delivery. • Create and maintain infrastructure-as-code templates using Terraform or other HashiCorp tooling. • Automate deployments, environment provisioning, configuration management, and release processes. • Ensure observability, logging, monitoring, and alerting are implemented and continuously improved. • Work closely with development teams to shape CI/CD workflows, build standards, and deployment strategies. • Enhance system reliability, scalability, and security across all environments. • Maintain cloud-native environments and ensure best practices for cost management, performance, and resilience. • Support incident response, troubleshooting, and root cause analysis for deployment- and environment-related issues. • Provide guidance and mentorship to development teams regarding DevOps practices, deployment patterns, and cloud-native architecture.
Advisor, Configuration – Release Engineering
MerativeA data and software partner for health and government social services, with tech and expertise to drive real progress.
• Develop and maintain the overall release strategy, including timelines, milestones, and deliverables for all Micromedex software releases (e.g. CMS, online, on premise, mobile) contributing to Micromedex business objectives. • Coordinate release schedules, cycles and activities operating across multiple disciplines (development, QA, operations, product management, support) to ensure timely delivery of software. • Manage and document release plans, including content, timelines, metrics and dependencies ensuring transparency for all stakeholders. • Manage maintenance backlog managing complexity to ensure prioritization of issues, effective triage, release planning, transparency, and clear communication on all issues raised. • Collaborate with support, product, engineering, operations and QA teams to ensure that all release criteria are met, including functionality, performance, compliance, security and deployment metrics achieved. • Identify, document, and mitigate risks that could impact the release schedule or quality. This includes contingency planning for delays or issues that arise during the release process. • Regularly communicate with stakeholders, including senior management, to provide updates on release status, support, risks, and outcomes. • Play your part in incident management ensuring timely response, evaluating mitigation options and prioritization. • Oversee post-release evaluations to gather feedback and lessons learned, facilitating continuous improvement in the release process. • Monitor and maintain deployment environments, ensuring their readiness and compatibility with upcoming releases. • Act as the primary point of contact during release activities, coordinating efforts to resolve issues and ensuring successful outcomes • Partner with leadership in transforming the engineering organization leading projects to deliver desired outcomes.
Cloud DevOps Engineer
Seamless Migration LLCDeveloper nerds who enable organizations through automation.
• Work in a highly critical production environment where automation and systems engineering expertise is essential • Support all three leading hyperscalers: AWS, Azure, and GCP • Develop and maintain solutions across multiple cloud platforms • Leverage infrastructure-as-code, CI/CD pipelines, and container orchestration tools to ensure reliable, scalable, and secure cloud operations • Develop and maintain scripts and automation tools to streamline deployment, configuration management, and system maintenance • Emphasize cloud-agnostic automation using tools like Terraform to build reusable infrastructure that works consistently across AWS, Azure, and GCP by leveraging common features across cloud providers • Implement efficient engineering and operational processes and establish best practices to continuously enhance operational efficiency within and across teams • Take full responsibility and end-to-end ownership of systems and tasks in production environments • Develop and maintain reusable Terraform scripts; modify and run infrastructure code from existing modules; troubleshoot deployment issues across AWS, Azure, and GCP environments • Develop new Ansible playbooks; modify, troubleshoot, and maintain existing Ansible code to automate infrastructure configuration and management tasks • Develop and maintain Python scripts to automate operational tasks, including AWS Lambda functions and Azure Functions, to support infrastructure automation • Design, implement, and maintain CI/CD pipelines using tools like GitLab CI and ArgoCD • Write and optimize pipeline scripts for automated testing and deployment • Leverage GitOps practices for managing infrastructure and application delivery




