Latitude.sh logo
Latitude.sh

Latitude.sh is a global bare metal cloud platform built for developers.

Senior Site Reliability Engineer

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 51-200Since 2001H1B No SponsorCompany SiteLinkedIn

Location

United States

Posted

131 days ago

Salary

0

Seniority

Senior

Job Description

Senior Site Reliability Engineer

Latitude.sh

• Continuously improve Latitude.sh’s platform reliability and performance • Design, build, and maintain tools to automate operational tasks and incident response • Implement and improve observability solutions, including monitoring, alerting, and tracing • Collaborate with engineering and platform teams to design scalable and resilient systems • Participate in on-call rotations and lead post-incident reviews with a focus on learning • Develop and document processes and runbooks that ensure operational excellence • Contribute to SLOs/SLIs definition and reliability metrics adoption across teams

Job Requirements

  • Strong verbal and written English communication skills
  • Advanced knowledge of Linux/Unix systems in production environments
  • Experience with Kubernetes and container orchestration
  • Proficiency with infrastructure automation tools (e.g., Terraform, Ansible)
  • Experience with observability stacks (e.g., Prometheus, Grafana, Loki, ELK)
  • Familiarity with scripting and programming languages such as Bash, Python, Go, or Ruby
  • Working knowledge of Git and CI/CD pipelines
  • Solid understanding of incident management and root cause analysis processes
  • Knowledge of cloud-native reliability and security best practices

Benefits

  • Paid Time Off
  • Competitive Compensation
  • Annual Bonus based on company and team performance
  • Flexible work hours
  • Opportunities for professional growth and development

Related Categories

Related Job Pages

More DevOps Engineer Jobs

NaNLABS logo

DevOps Engineer

NaNLABS

Your Sidekick for AI, Cloud-Native & Real-Time Data Engineering | Scalable Innovations in Auto, EV, SaaS & Cybersecurity

DevOps Engineer131 days ago
Full TimeRemoteTeam 51-200Since 2013H1B No Sponsor

• Design, implement, and maintain scalable and secure cloud infrastructure in AWS • Manage infrastructure as code using Terraform to ensure consistency, automation, and reliability across environments • Build and improve CI/CD pipelines using GitHub Actions to enable efficient and reliable software delivery • Manage Kubernetes infrastructure (EKS) and support networking, security, and access management within AWS environments • Implement observability practices across services using modern monitoring and alerting tools • Support and scale infrastructure for data and machine learning workloads • Collaborate closely with engineering and data teams to ensure infrastructure supports product scalability and performance • Define and promote DevOps best practices related to automation, reliability, and infrastructure standards • Participate in incident response and help improve reliability through proactive monitoring and infrastructure improvements • Communicate technical decisions, risks, and trade-offs clearly with both technical and non-technical stakeholders

Latin America
Job Closed
Nearsure logo

Senior DevOps Engineer

Nearsure

Remove the barriers to growth by scaling your team fast with top-notch Latin American IT talent

DevOps Engineer131 days ago
Full TimeRemoteTeam 201-500H1B No Sponsor

• Design, build, and standardize CI/CD pipelines (Azure DevOps preferred) to support multiple development teams. • Build and evolve a self-service, GitOps-driven platform on Azure. • Contribute to the architecture, design, and implementation of Azure Kubernetes Service (AKS) environments. • Contribute to architecture and security configuration decisions within Azure and Kubernetes. • Implement Infrastructure as Code using Terraform and Helm to standardize and automate cloud environments. • Establish reusable automation patterns and platform standards across teams. • Integrate and automate security tooling within CI/CD pipelines in alignment with the dedicated security team. • Maintain and optimize existing pipelines and platform components. • Standardize CI/CD and deployment practices across 7–8 parallel teams. • Enable development teams through platform improvements and self-service capabilities. • Troubleshoot and debug production issues across services, infrastructure, and Kubernetes clusters. • Participate in incident response and post-mortem activities when required. • Collaborate closely with cross-functional and distributed teams to ensure consistent platform evolution. • Take ownership of platform components, driving improvements proactively and independently.

Latin America
Job Closed
Flex Dental Solutions logo

Site Reliability Engineer

Flex Dental Solutions

Flex is a collection of smart and easy-to-use tools that pair with Open Dental to supercharge your patient engagement.

DevOps Engineer131 days ago
Full TimeRemoteTeam 11-50H1B No Sponsor

• Own the availability, performance, and resilience of our production systems. • Partner closely with engineering, product, and leadership to reduce operational risk. • Proactively monitor application health and performance across cloud infrastructure (AWS). • Lead incident response, including triage, mitigation, root cause analysis (RCA). • Lead and participate in disaster recovery drills and security incident simulations. • Build and maintain Infrastructure as Code (IaC) using AWS-native tooling. • Collaborate with development teams to improve CI/CD reliability. • Work closely with stakeholders and product teams to ensure technical reliability aligns with business needs. • Reduce operational toil through automation, tooling, and process improvements. • Champion best practices across security, availability, performance, and incident response.

United States
Job Closed
Full TimeRemoteTeam 501-1,000H1B No Sponsor

• Diseñar y mantener plataformas de autoservicio y marcos de automatización para reducir fricciones en la entrega de software. • Optimizar pipelines de CI/CD utilizando herramientas como GitHub Actions, ArgoCD, Argo Rollouts y Argo Workflows para implementaciones escalables, seguras y eficientes. • Automatizar el aprovisionamiento de infraestructura y la gestión de configuración con herramientas como Terraform, Pulumi, Crossplane o ACK, adoptando un enfoque GitOps. • Diseñar soluciones en la nube en AWS, siguiendo las mejores prácticas en seguridad, observabilidad y optimización de costos. • Monitorear y optimizar el rendimiento de la plataforma utilizando herramientas como Grafana, Datadog, New Relic, Coralogix y Prometheus. • Participar en proyectos de migración hacia AWS, asegurando una transición eficiente y optimizada.

Argentina