Site Reliability Engineer

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 1,001-5,000H1B No SponsorCompany SiteLinkedIn

Location

Costa Rica

Posted

11 days ago

Salary

0

Seniority

Senior

Job Description

Site Reliability Engineer

CSC Generation

• Work on service resiliency, performance tuning, and system design across Backcountry's platform • Drive resolution of critical incidents and ensure fixes are methodically implemented through postmortems • Leverage AI-assisted engineering tools (Claude Code, GitHub Copilot, MCP-based agents) to investigate, automate, and ship fixes across infrastructure and application repositories • Reduce toil by designing and implementing automation • Partner with other Site Reliability Engineers, developers, and architects to evaluate and implement best practices for current and future workloads • Monitor system health and capacity, taking proactive action to fix problems before they occur • Collaborate with engineering teams to build, deploy, and support features • Build and maintain observability (metrics, logs, traces, profiles) and SLI/SLO instrumentation for Backcountry services • Participate in FinOps initiatives across GCP and AWS, including capacity planning and committed-use discount strategy • Participate in the on-call support rotation within the SRE team

Job Requirements

  • 3+ years of experience supporting containerized production services, preferably running Kubernetes
  • 3+ years of experience with Infrastructure as Code (Terraform, AWS CDK, Ansible, etc.)
  • 3+ years of cloud experience operating in Google Cloud Platform and/or AWS (multi-cloud stack; Azure/Entra exposure is a plus)
  • Comfortable diagnosing issues and shipping bug fixes directly to application code (not just infrastructure) to keep services reliable and stable
  • Comfortable performing deep dives across both infrastructure and application/software git repositories to trace issues end-to-end
  • Proficient with AI-assisted coding tools (e.g., Claude Code, GitHub Copilot) and MCP-based agents, used to accelerate investigation, code review, and automation
  • Strong knowledge of scripting and programming languages (Bash, Python, and TypeScript/Node.js)
  • Experience managing Linux (any major distribution) in production environments
  • Excellent understanding of internet application protocols (DHCP, DNS, HTTPS, SSH, etc.)
  • Understanding of how DevOps (CI/CD) and SRE practices (SLOs, SLIs) apply to daily work
  • Hands-on experience with observability tooling (Grafana, Prometheus, Loki, OpenSearch, or equivalents) and SLI/SLO instrumentation
  • Experience with GitOps and Kubernetes packaging (ArgoCD, Helm, Kustomize)
  • Proactively track emerging technology trends and developments, evaluating which ones are worth bringing into engineering practice
  • Bachelor's degree in computer science or similar, or equivalent experience
  • Advanced-level English communication skills, both verbal and written.

Benefits

  • Competitive Benefits: We offer an attractive benefits package including primarily remote work, private medical and life insurance, additional paid time off, monthly allowances and reimbursements, employee discounts, and opportunities for professional growth.

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Sólides logo

Senior Site Reliability Engineer

Sólides

Com a Sólides seu time Joga Fácil ⚽

DevOps Engineer11 days ago
Full TimeRemoteTeam 501-1,000

• Monitoring and Alerts: Build, maintain and evolve clear dashboards and intelligent alerts (infrastructure and business rules), ensuring real-time visibility into system health; • Log Management and APM: Actively analyze, parse and centralize logs, and configure and monitor APM metrics to optimize performance; • Automation: Develop and maintain automation solutions for provisioning, configuration and deployment of infrastructure (IaC); • Incident Management: Participate in resolving production problems and incidents, using observability data for rapid diagnostics and root cause analysis; • SRE Culture: Collaborate with and promote best practices among development teams for resilience, instrumentation and metrics collection.

Brazil
Sólides logo

Senior DevOps Engineer

Sólides

Juntos com quem sonha grande 💜

DevOps Engineer11 days ago
Full TimeRemoteTeam 501-1,000H1B No Sponsor

• CI/CD Pipelines (Continuous Integration and Continuous Delivery): Build, maintain and evolve automated pipelines for building, testing and deployment, ensuring fast, secure and frequent releases • Infrastructure as Code (IaC): Provision, manage and evolve cloud infrastructure through versioned code, ensuring standardization and consistency across environments • Orchestration and Containers: Design, maintain and optimize the container ecosystem, defining efficient deployment strategies (e.g., Blue-Green, Canary) and high availability • Security and DevSecOps: Integrate security practices and checks throughout the development lifecycle (shift-left), managing secrets, access policies and automated vulnerability analysis such as SAST and DAST • DevOps Culture and Collaboration: Collaborate closely with development teams to remove operational bottlenecks, promoting autonomy, agility and knowledge sharing across engineering teams.

Brazil
Full TimeRemoteTeam 201-500H1B Sponsor

• Build the Future of Cloud Delivery • Shape and scale our DevOps practice. • Lead & Grow a Team: Mentor and develop a high-performing DevOps team (2–5 engineers), fostering collaboration, ownership, and continuous improvement. • Own the DevOps Strategy: Define and execute our Azure DevOps vision and roadmap aligned with business and engineering goals. • Stay Hands-On: Design, build, and maintain CI/CD pipelines using Azure DevOps and modern tooling. • Drive Cloud & Modernization Initiatives: Lead efforts to modernize applications and platforms in a cloud-first environment. • Enable Scalable Delivery: Partner with engineering teams to improve reliability, scalability, security, and speed. • Implement Infrastructure as Code: Use ARM, Bicep, Terraform, or similar tools to build and manage infrastructure. • Support Data & Analytics Pipelines: Build CI/CD for Azure Data Factory, Snowflake, Power BI, and database deployments. • Advance DevSecOps & MLOps: Implement secure, automated pipelines for applications, data, and AI/ML workloads. • Optimize & Govern: Improve cloud performance, cost efficiency, and ensure compliance and best practices. • Act as a Technical Leader: Serve as an escalation point and collaborate across architecture, security, and product teams.

New Hampshire
Job Closed
Fundraise Up logo

Senior DevOps Engineer

Fundraise Up

Unlocking the world's generosity potential

DevOps Engineer11 days ago
Full TimeRemoteTeam 51-200H1B No Sponsor

• Own one of our core platform areas end-to-end: observability (VictoriaMetrics, Grafana, Graylog / VictoriaLogs, fluent bit, exporters, alerting) or CI/CD (Jenkins scripted pipelines, Harbor, Nexus, build agents) — you drive its architecture, reliability, and roadmap. • Drive technical initiatives end-to-end: gather requirements, write the design doc, decompose into tasks, implement, deliver to production, and own the operational health afterwards. • Drive clarity in ambiguous situations by defining requirements, assumptions, and next steps. • Design for reliability and scale: evolve the architecture of our platforms — topology, integration points, scaling approach, and reliability model. • Support developers: deploy and monitor applications on both on-premise servers and Kubernetes (Helm), troubleshoot builds and deploys, help teams with metrics, alerts, and logs; participate in chat duty in developer support channels. • Automate away toil: repetitive operations, provisioning, and maintenance should be codified, not performed by hand. • Investigate production incidents as the senior escalation point for your area: drive resolution, lead post-mortems, implement systemic fixes. Participate in on-call rotations and raise the bar for how on-call works. • Mentor less experienced engineers through design discussions, reviews, and pairing; catch debt-inducing shortcuts at the review stage. • Use AI in all aspects of day-to-day work: researching, troubleshooting, developing.

Georgia
₾14.3K - ₾16.0K / month