Build software faster. The One DevOps Platform enables your entire org to collaborate around your code. We're hiring.
Site Reliability Engineer
Location
Worldwide
Posted
10 days ago
Salary
0
Seniority
Mid Level
No structured requirement data.
Job Description
Site Reliability Engineer
GitLab
Role Description We're looking for an Intermediate Site Reliability Engineer to join the Runners Platform team. In this role, you'll build and operate Hosted Runners for GitLab Dedicated, the managed CI/CD compute platform that runs our customers' pipelines inside single-tenant Amazon Web Services (AWS) environments. You'll own infrastructure automation across the full lifecycle: - Provisioning runner fleets with Terraform - Extending the Go tooling and autoscaling stack - Ensuring customer continuous integration and continuous delivery (CI/CD) jobs run reliably across the platform What you’ll do: - Design, build, and operate AWS infrastructure for Hosted Runners across many single-tenant environments, including: - Elastic Compute Cloud (EC2) - Auto Scaling Groups - Virtual Private Cloud (VPC) networking - Subnets - Network Address Translation (NAT) - Network access control lists - PrivateLink - Identity and Access Management (IAM) - Elastic Container Registry (ECR) - Develop and maintain infrastructure as code using Terraform, contributing to common modules and the deployment tooling that provisions and upgrades runner stacks. - Write Go code for our runner tooling and autoscaling components, including: - Fleeting instance-lifecycle plugins - Zero-downtime deployment command-line interface - Reusable infrastructure toolkits - Build and improve the GitLab CI/CD pipelines that orchestrate: - Blue/green zero-downtime deployments - Automated upgrades - Quality assurance validation - Performance testing of runner stacks - Define and monitor service level objectives for CI job execution, including: - Queue times - Job success rates - Fleet saturation - Build the Grafana dashboards, alerts, and runbooks behind service level objectives. - Participate in an on-call rotation, handle incidents affecting customer CI/CD workloads, and automate away recurring toil. - Run performance and scale testing that reflects real customer workloads, and tune autoscaling parameters for cost and reliability. - Write documentation and runbooks for consistent operation of runner stacks. Qualifications - Professional experience operating production infrastructure on AWS at scale, including EC2, Auto Scaling Groups, IAM, and VPC networking. - Strong infrastructure-as-code experience with Terraform, including writing and refactoring modules used by other teams. - Proficiency in Go for building and debugging infrastructure tooling, or strong experience in another systems language and willingness to work in Go daily. - Practical knowledge of CI/CD systems and job execution. - Experience with observability practices such as metrics, dashboards, alerting, logging, and service-level-objective-based monitoring. - Experience with on-call rotations and incident management for customer-facing systems. - Strong problem-solving skills, excellent written communication, and comfort working asynchronously across multiple time zones. - Direct GitLab Runner experience, familiarity with configuration management such as Ansible, and container tooling such as Docker are a plus. Benefits - Flexible Paid Time Off - Team Member Resource Groups - Equity Compensation & Employee Stock Purchase Plan - Growth and Development Fund - Parental Leave
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Senior DevSecOps Engineer
Bet On TalentYour go-to place for career opportunities in iGaming, Gaming and Fintech.
• Set up, implement, and support the technological solutions that keep our systems running smoothly and securely. • Build and maintain CI/CD pipelines that power our production deploys, keeping releases fast and reliable. • Develop and continuously improve our monitoring and alerting systems, so issues get caught before they become problems. • Create and maintain "infrastructure as code" configurations to manage our cloud environments on AWS and Cloudflare. • Implement methodologies and procedures that make our systems resilient, capable of recovering automatically from failure with minimal to no manual intervention. • Act as a key point of communication across departments, bridging engineering with the rest of the organization.
• Deploy servers with hardening and firewall policies between servers • Support the deployment of applications developed by Azeus in the cloud (AWS, Azure, KSA cloud, etc.) • Carry out the deployment of applications developed by Azeus on client premises. You will be prepared to handle server and application software hardening, firewall policy setup, etc. • Identify and analyze system integration problems • Resolve server issues by identifying the root cause and taking immediate remedial action • Formulate and implement corrective actions in case a problem re-occurs • Formulate and implement preventive actions to prevent issues from happening again • Identify recurring tasks and recommend tools to automate those tasks • Be on-call for support cases
Team Lead Cloud & DevOps
ARES Consulting GmbHThe Cloud Native Company: Experten und Teams für die Bereiche Cloud Native Development, Cloud Admin und DevOps
Role Description Als Team Lead (m/w/d) übernimmst du strategische Verantwortung und bleibst gleichzeitig noch nah an der Technik – mit Fokus auf die Transformation hin zu sicheren, skalierbaren und zukunftsfähigen Cloud-Infrastrukturen, insbesondere für Kunden aus dem öffentlichen Sektor. - Fachliche und disziplinarische Verantwortung für ein erfahrenes Team aus Cloud & DevOps Engineers, inklusive technischer Steuerung von Projektteams, Projektkoordination und Ressourcenplanung - Aktive Weiterentwicklung deiner Teammitglieder – fachlich wie persönlich - Ungefähr 60% deiner Zeit bringst du direkt in Kundenprojekte ein – durch Architektur und Konzeption mit Fokus auf Private-Cloud-Umgebungen sowie ausgewählten Public-Cloud-Plattformen (z. B. AWS, Azure, STACKIT) - Product Ownership und Entwicklung technischer Visionen für unsere Kunden - Unterstützung bei Ausschreibungen und Angeboten sowie technische Unterstützung in Presales-Gesprächen - Strategische Mitgestaltung des Wachstums von ARES Qualifications - Abgeschlossenes Studium oder Ausbildung im IT-Umfeld - Einschlägige Branchen-Erfahrung im öffentlichen Dienst oder KRITIS-Umfeld - Erfahrung in der Auswahl und disziplinarischen Führung von Personal - Erfahrung in der Konzeption und Architektur moderner Cloud- und Plattformlösungen, insbesondere in Private-Cloud- oder hybriden Umgebungen - Fundierte Kenntnisse des Kubernetes-Ökosystems - Verständnis von Public Cloud-Providern und deren Services wünschenswert (z.B. AWS, Azure, GCP, STACKIT) - Erfahrungen im Presales oder der technischen Kundenberatung sind von Vorteil - Du kannst technische Entscheidungen souverän begründen - IT-Security denkst du in allen technischen Entscheidungen mit - Selbstständige, kommunikative Arbeitsweise mit ausgeprägter Kundenorientierung - Sehr gute Deutsch- und Englischkenntnisse - Reisebereitschaft innerhalb Deutschlands (im Schnitt max. 3 Tage/Monat) - Bereitschaft zur Sicherheitsüberprüfung (SÜ1 oder SÜ2) nach SÜG Benefits - Flexibilität & Remote-First: Flexible Arbeitszeitgestaltung und 100% Homeoffice neben gelegentlicher Reisetätigkeit (innerhalb Deutschlands) - Moderne Hardware & Arbeitsausstattung: MacBook oder Windows nach Wahl und zusätzliches Equipment für dein Homeoffice-Setup - Urlaub & Bonus: 30 Tage Urlaub/Jahr sowie leistungsabhängiger Jahresbonus (bis zu 1 Monatsgehalt) - Weiterbildung: Jährliches Weiterbildungsbudget sowie individuelle Karriereplanung und Mentoring - Team & Kultur: Umfassendes Onboarding- und Buddy-Programm, mehrere Teamevents pro Jahr sowie kurze Kommunikationswege und ein starkes Team, das Wissen teilt und gemeinsam etwas bewegt - Corporate Benefits: Wähle zwischen Urban Sports Club, Wellpass oder einer SpenditCard im Wert von 50 EUR monatlich (steuerfrei) und profitiere zusätzlich von Angeboten wie Firmenwagen, D-Ticket Zuschuss, betrieblicher Altersvorsorge und Mitarbeiterrabatten Company Description Als cloud-native Company beraten und unterstützen wir öffentliche Auftraggeber und regulierte Organisationen hands-on in ihren Projekten mit Fokus auf moderne Cloud-, Plattform- und Digital-Workplace-Lösungen. Unsere Leistungen reichen dabei von der strategischen Konzeption über die Implementierung bis hin zum operativen Betrieb. Der Schlüssel zu unserem Erfolg? Unser starkes Team, welches auf höchstem, technischen Niveau agiert und eine New-Work-Culture, die wir nicht nur versprechen, sondern auch jeden Tag leben. Werde Teil von ARES und gestalte aktiv mit uns die digitale Zukunft Deutschlands!
DevOps Engineer – Regular/Senior
SpyrosoftWe enable our clients to thrive, thanks to a combination of technical proficiency and domain-specific knowledge.
• Manage, maintain, and optimize OpenShift/Kubernetes environments. • Develop and improve GitOps-based deployment processes using ArgoCD. • Design, build, and maintain CI/CD pipelines using Bamboo, Bitbucket, and related tools. • Ensure platform reliability, performance, and high availability. • Implement and enhance monitoring, logging, and alerting solutions using Prometheus, Grafana, Elastic Stack, and Logstash. • Automate operational tasks and infrastructure management using scripting and DevOps best practices. • Support application deployments across multiple environments. • Manage secrets and security-related configurations using Vault and enterprise security standards. • Troubleshoot production issues and participate in incident resolution activities. • Collaborate closely with development, QA, and operations teams to improve delivery processes and system stability. • Participate in shift work and scheduled on-call support rotations. • Contribute to continuous improvement initiatives focused on automation, observability, security, and operational excellence.




