Job Closed

This listing is no longer active.

Galaxy logo
Galaxy

Engineering a new economic paradigm.

Vice President – Site Reliability Engineering, Data Centers

DevOps EngineerDevOps EngineerFull TimeRemoteLeadTeam 201-500Since 2018H1B SponsorCompany SiteLinkedIn

Location

United States

Posted

71 days ago

Salary

0

Seniority

Lead

Job Description

Vice President – Site Reliability Engineering, Data Centers

Galaxy

• Oversee a specialized SRE team focused on the design, deployment, and maintenance of automation toolsets as well as the systems they interact with. • Establish and enforce standards for IaC to ensure consistent, repeatable, and secure deployments across an entire infrastructure ecosystem. Strong proficiency in Terraform is required. • Lead the strategy for automated configuration and state management, ensuring Ansible playbooks and Packer image pipelines are optimized for both Windows, Linux, and ESXi Platforms. • Manage the monitoring and health of the automation platforms themselves. Implement SLIs/SLOs to ensure the 'tools that build the servers' are highly available and performant. • Drive the automated lifecycle of both physical and virtual assets, from initial template creation/deployment to automated patching, scaling, and decommissioning. • Lead the development of custom scripts and internal providers (Python, Go, PowerShell, Bash) to provide better insights and tooling for our systems. • Outside of the automation team you will need to be able to collaborate and foster workflows alongside the rest of the Datacenter team and be able to facilitate needs for the team as a whole. • Analyze system behavior and resource utilization in virtual environments to optimize the performance of automated deployments. • Provide technical guidance and career mentorship to SREs, fostering a culture of 'automate-first' and continuous improvement.

Job Requirements

  • 6-10 years’ experience in Infrastructure, SRE or DevOps, specifically focused on infrastructure automation at scale.
  • Deep proficiency with Terraform (providers, modules, state management) and Ansible (roles, playbooks, Tower/AWX).
  • Hands-on experience with Image Creation (i.e. Packer, Ansible, SCCM) to build standardized, hardened images for both Windows and Linux in hybrid environments.
  • Strong experience managing and automating virtual platforms such as VMware (vSphere/vCenter) as well as Cloud providers such as Azure and AWS.
  • High-level scripting skills in mediums such as Python, Go, PowerShell, and Bash.
  • Experience with observability tools (Splunk, ELK, Prometheus, or Grafana) to monitor infrastructure health and automation telemetry.
  • Good understanding of Network topology and design as well as experience with platforms such as Juniper Networks or Palo Alto.
  • Strong mastery of Git (branching strategies, PR workflows) and CI/CD platforms (Jenkins, GitLab CI, or GitHub Actions).
  • Equal comfort managing, troubleshooting, and tuning performance for both Windows Server and Linux.

Benefits

  • Flexible work arrangements
  • Professional development opportunities

Related Categories

Related Job Pages

More DevOps Engineer Jobs

viind GmbH logo

DevOps Engineer

viind GmbH

Chatbots für Verwaltungen und Privatunternehmen | Systeme mit KI-Anbindung | Recruiting via Messenger

DevOps Engineer71 days ago
Full TimeRemoteTeam 11-50H1B No Sponsor

• You take responsibility for the operation, maintenance and further development of our Kubernetes clusters (Hetzner Cloud & on‑premises) • You ensure the availability, scalability and security of our infrastructure and continuously optimize it • You operate and maintain our central platform components such as databases (PostgreSQL, Typesense, MongoDB), Keycloak and self‑hosted AI models • You develop and implement strategies for deployments, updates, backups and recovery of these systems • You implement and run monitoring and logging solutions for our entire infrastructure • You develop, operate and optimize our CI/CD pipelines based on GitLab and Docker • You contribute your own ideas to continuously improve and evolve our infrastructure, processes and tools • You manage our internal cloud and SaaS services (e.g. Atlassian, Microsoft 365, GitLab) • You ensure that all systems and processes meet the requirements of ISO 27001 and GDPR and actively support audits • You work closely with our development team and support infrastructure-related questions or backend development

Germany
Job Closed
Cribl logo

Senior Site Reliability Engineer

Cribl

Cribl, the Data Engine for IT and Security, empowers organizations to transform their data strategy.

DevOps Engineer71 days ago
Full TimeRemoteTeam 501-1,000Since 2017H1B Sponsor

• Engage with teams and improve service delivery and reliability across their entire lifecycle • Measure and monitor all production systems with an eye towards availability, latency and overall system health • Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence • Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability • Help identify and drive down toil with creative innovation and automation • This position will require stand-by, on-call, or off-hours duties

Poland
Resilient Co. logo

Senior DevOps

Resilient Co.

WE ARE RESILIENT CO. We adapt to your needs.

DevOps Engineer71 days ago
ContractRemoteTeam 11-50Since 2020H1B No Sponsor

• Design and implement infrastructure-as-code using Terraform for Azure services including AKS, Blob Storage and App Services. • Build, maintain and optimize CI/CD pipelines and mobile/web build pipelines. • Operate, troubleshoot and tune Kubernetes and Docker-based workloads running on AKS. • Implement and manage SSO and External ID flows using Microsoft Entra. • Create reusable templates, Terraform modules and pipeline templates to enable developer self-service. • Collaborate directly with technical leads to define platform direction and deployment patterns. • Mentor engineers on deployment best practices, observability and platform usage. • Own platform-level decisions and improvements, prioritizing strategic work over ticket-level execution. • Write clear, async-friendly documentation and communicate effectively in AI-augmented workflows. • Manage and support PostgreSQL-related deployment and operational concerns as they relate to platform infrastructure.

Argentina
Job Closed
SupplyHouse.com logo

Site Reliability Engineer

SupplyHouse.com

Plumbing, Heating & HVAC Supplies. Real People. Real Service.

DevOps Engineer71 days ago
Full TimeRemoteTeam 501-1,000Since 2004H1B Sponsor

• Design, build, and maintain scalable, reliable systems on GCP (Compute Engine, GKE, Cloud Storage, Cloud SQL) • Develop automation for infrastructure provisioning using Terraform, Ansible, or Deployment Manager • Build and maintain observability platforms (monitoring, logging, tracing) using tools such as Stackdriver (Cloud Monitoring), Prometheus, or Grafana • Manage incident response, conduct postmortems, and implement improvements to reduce recurrence • Partner with DevOps and engineering teams to enhance CI/CD pipelines for resilient deployments • Define and monitor SLAs, SLOs, and SLIs to ensure application availability and performance • Implement disaster recovery (DR) and backup strategies across cloud services • Continuously optimize performance, capacity, and cost-efficiency of GCP resources

India
$29K - $36K / year
Job Closed