Group-IB logo
Group-IB

Global Threat Hunting and Adversary-Centric Cyber Intelligence Company

Senior DevOps Engineer

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 501-1,000Since 2003Company SiteLinkedIn

Location

Spain

Posted

5 days ago

Salary

0

Seniority

Senior

Job Description

Senior DevOps Engineer

Group-IB

• Maintain and optimize core server infrastructure, including bare-metal servers, LXC containers, virtual machines, and cloud environments. • Operate and support core infrastructure services such as Nginx, Puppet, GitLab, Artifactory, Nexus, Harbor, Grafana, etc. • Manage and evolve infrastructure following the Infrastructure as Code (IaC) paradigm. • Handle and resolve incidents related to infrastructure operations. • Collaborate closely with cross-functional teams (network engineers, developers, and other technical stakeholders). • Design and implement high-availability, fault-tolerant, and scalable software solutions. • Monitor service performance and availability using modern observability tools, ensuring system reliability and optimal resource utilization.

Job Requirements

  • 4+ years of experience in Linux administration, DevOps, or Site Reliability Engineering (SRE).
  • Strong proficiency in automating tasks using Bash or similar scripting languages.
  • Solid understanding of networking fundamentals (TCP/IP stack, routing, DNS, etc.).
  • Hands-on experience managing bare-metal infrastructure in production environments.
  • Experience with configuration management systems such as Ansible or Puppet.
  • Experience with distributed databases (Elasticsearch, Cassandra, MongoDB, MySQL, PostgreSQL, etc.).
  • Expertise in Kubernetes administration and managing containerized workloads.
  • Experience with IaC tools such as FluxCD or ArgoCD.
  • Ability to design and implement high-performance, fault-tolerant, and secure infrastructure solutions.
  • Experience with monitoring and observability systems — Zabbix, VictoriaMetrics, Loki, Grafana, etc. — including building dashboards and configuring alerting.

Benefits

  • Work with real stakes. Group-IB investigates active cybercriminal groups, responds to breaches affecting critical infrastructure, and develops technologies used by law enforcement agencies including INTERPOL, Europol, and Afripol across 60+ countries. We've conducted 1,550+ cybercrime investigations alongside 600+ enterprise customers globally. When you join Group-IB, your work directly disrupts digital crime.
  • Grow your way. Choose your own path: deepen your craft as a technical expert, step into leadership, move across to another team, or relocate to one of our Digital Crime Resistance Centers across the Americas, Europe, the Middle East & Africa, Central Asia, and the Asia-Pacific. Your growth is our growth — Group-IB's expansion across 60+ active country operations means real career acceleration.
  • We fund professional certifications at company expense — whether you're pursuing CEH, CISSP, OSCP, or specialized certifications in forensics and penetration testing. You don't have to choose between doing the job and advancing your credentials.
  • Work alongside industry leaders. Our Unified Risk Platform — Threat Intelligence, Digital Risk Protection, Attack Surface Management, Managed XDR, and more — is recognized by Gartner, Forrester, KuppingerCole, and Datos Insights. Frost & Sullivan named us a 2025 Global Technology Innovation Leader. When you work here, you're building technologies that set the industry standard.
  • Real challenges, real expertise. You'll take on complex, real-world problems alongside adversary-centric researchers and incident response experts spread across six continents. We've built 21+ years of proprietary telemetry through 1,500+ joint investigations. No two threats look alike — and neither do the skills you'll develop.
  • A team that is genuinely international. Our people come from different countries, speak different languages, and bring different perspectives. What connects us is a shared mission: fighting cybercrime and making the world safer. We care about your wellbeing and happiness as much as your output.

Related Categories

Related Job Pages

More DevOps Engineer Jobs

IT42morrow IFT GmbH logo

DevOps Engineer

IT42morrow IFT GmbH

Mit IT42morrow heute die IT-Lösungen von morgen entwickeln.

DevOps Engineer5 days ago
Full TimeRemoteTeam 1-10H1B No Sponsor

Role Description IT42morrow steht für Innovation und Dynamik in der IT-Welt. Mit einer Palette von spannenden Projekten und einer Kultur, die Wachstum und Diversität fördert, sind wir auf der Suche nach einem neuen Star in unserem DevOps-Team. Egal, ob du ein erfahrener Profi oder ein begeisterter Neuling mit einer Leidenschaft für Technologie bist – wenn du bereit bist, in die Welt der großen IT-Umgebungen einzutauchen, sollten wir uns unterhalten! - Nimm das Ruder in die Hand bei der Betreuung von Cloud-Umgebungen und der Steuerung von CI/CD-Pipelines. - Zeige deine Skills in der Automatisierung und Optimierung von Prozessen. - Werde zum Detektiv, wenn es darum geht, Schwachstellen aufzuspüren. - Du baust Brücken zwischen Entwicklung und Infrastruktur? Das ist dein Ding! Qualifications - Tech-Enthusiast: Neue Technologien ziehen Sie magisch an. Du bist immer up-to-date und brennen darauf, Innovationen praktisch anzuwenden. - CI/CD: Jenkins, Gitlab, Azure DevOps, Bitbucket - Cloud: Amazon Webservices, Microsoft Azure, Google Cloud Platform - Automatisierung: Ansible, Puppet, Chef, SaltStack - Container(-Orchestrierung): Docker, Kubernetes, OpenShift, Rancher - Monitoring/Observability: Prometheus, Nagios, ELK-Stack, Grafana - Entwicklung: Java, Python, Go, Golang - Sicherheitsbewusst: Für dich steht die Sicherheit der Systeme immer an erster Stelle. Du kennst die Best Practices und sorgen dafür, dass Sicherheitslücken keine Chance haben. - Kommunikationsstark: Du weißt, dass gute Ideen geteilt werden müssen. Im Team fühlst du dich wohl und kannst dein Wissen klar und verständlich vermitteln. - Lernwillig: Die IT-Welt verändert sich rasant, und du bist immer bereit, Neues zu lernen und sich weiterzuentwickeln. - Reiselustig: Du bist offen für Einsätze bei unseren Kunden und sehen darin eine Chance, deinen Horizont zu erweitern. - Sprachtalent: Du beherrschst Deutsch mindestens auf dem Niveau C1. Englischkenntnisse sind ein absoluter Pluspunkt. Benefits - Zukunftssicherheit: Freuen Sie sich auf einen unbefristeten Arbeitsplatz in einem zukunftsorientierten Unternehmen. - Entwicklungspotenzial: Wir investieren in Ihre Karriere mit individuellen Weiterbildungsmöglichkeiten. - Cutting-Edge Technologie: Arbeiten Sie mit den neuesten Technologien und bleiben Sie immer am Puls der Zeit. - Projektvielfalt: Bei uns gibt es eine breite Palette an Herausforderungen, die Ihren Arbeitsalltag spannend und abwechslungsreich gestalten. - Reisebereitschaft: Für Dienstreisen sind Sie bei uns in guten Händen – Organisation und Kostenübernahme sind selbstverständlich. Company Description Große IT-Umgebungen, spannende Kunden und Projekte, IT für morgen – das ist IT42morrow! Wir wachsen – uns das mit dir. Unsere Kunden unterstützen wir in ganz Deutschland in Projekten in der Software-Entwicklung, im Bereich DevOps und im Software-Testing.

Germany
Full TimeRemoteTeam 11-50Since 2018

• Own the operational health of computer-vision / AI models and associated services. • Manage reliability and operations, defining SLIs and SLOs for key services. • Ensure observability is trustworthy, improving alert quality and building necessary metrics. • Lead incident response, driving recovery and producing actionable postmortems. • Plan capacity and performance to preempt saturation. • Own business continuity, disaster recovery, and backup strategies. • Engineer software solutions for operational efficiency and reliability. • Govern production change management collaboratively and safely. • Maintain CI/CD pipelines and Terraform configurations for the AWS estate. • Harden security and compliance within the system.

United States
Full TimeRemoteTeam 11-50Since 2018

• Own service objectives. Define SLIs and SLOs for the services that matter, manage error budgets, and use them to drive prioritization and change-rate decisions. Make reliability a measured number, not a feeling. • Make observability trustworthy. Own alert quality end-to-end: high signal-to-noise, alarms that reliably catch the incidents they are meant to catch, and the discipline to turn off a misleading alarm until it is fixed properly. Build the dashboards and the custom metrics, exporters, and instrumentation (CloudWatch, OpenTelemetry) needed to see the system clearly. • Lead incident response. Run incidents calmly, drive mean-time-to-recovery down, and produce blameless postmortems with action items that actually get closed. Improve and own the on-call rotation and its health. • Plan capacity and performance. Forecast and right-size compute (especially GPU), Kafka/MSK throughput and partitioning, RDS/TimescaleDB load, and Redis. Catch saturation before customers do. • Own business continuity and disaster recovery. Backups, replication, failover, and recovery for RDS, MSK, Redis, and the edge fleet. Define RPO/RTO and prove them with regular, tested game days, not assumptions. • Keep the edge fleet healthy. Remote diagnosis and recovery over AWS IoT, container auto-update over systemd timers, and the ongoing CentOS 7 migration to the containerized media stack (Ubuntu 22.04). • Engineer away toil. Write real software (Python, Golang, Bash) to automate operational work, self-heal common failures, and make reliability repeatable instead of heroic. • Govern production change safely. Enforce collaborative, reviewed change management; protect the system from risky, unilateral changes (topology, instance-count, scaling, and config).

United Kingdom
Full TimeRemoteTeam 11-50Since 2018

• Own service objectives. Define SLIs and SLOs for the services that matter, manage error budgets, and use them to drive prioritization and change-rate decisions. Make reliability a measured number, not a feeling. • Make observability trustworthy. Own alert quality end-to-end: high signal-to-noise, alarms that reliably catch the incidents they are meant to catch, and the discipline to turn off a misleading alarm until it is fixed properly. Build the dashboards and the custom metrics, exporters, and instrumentation (CloudWatch, OpenTelemetry) needed to see the system clearly. • Lead incident response. Run incidents calmly, drive mean-time-to-recovery down, and produce blameless postmortems with action items that actually get closed. Improve and own the on-call rotation and its health. • Plan capacity and performance. Forecast and right-size compute (especially GPU), Kafka/MSK throughput and partitioning, RDS/TimescaleDB load, and Redis. Catch saturation before customers do. • Own business continuity and disaster recovery. Backups, replication, failover, and recovery for RDS, MSK, Redis, and the edge fleet. Define RPO/RTO and prove them with regular, tested game days, not assumptions. • Keep the edge fleet healthy. Remote diagnosis and recovery over AWS IoT, container auto-update over systemd timers, and the ongoing CentOS 7 migration to the containerized media stack (Ubuntu 22.04). • Engineer away toil. Write real software (Python, Golang, Bash) to automate operational work, self-heal common failures, and make reliability repeatable instead of heroic. • Govern production change safely. Enforce collaborative, reviewed change management; protect the system from risky, unilateral changes (topology, instance-count, scaling, and config).

Serbia