Job Closed

This listing is no longer active.

Site Reliability Engineer

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 1,001-5,000H1B No SponsorCompany SiteLinkedIn

Location

Georgia

Posted

132 days ago

Salary

0

Seniority

Senior

Job Description

Site Reliability Engineer

Intermedia Cloud Communications

• Build and operate metrics/monitoring platforms: Prometheus and/or VictoriaMetrics (scrape configs, exporters, recording rules) • Design and maintain alerting strategy: thresholds, anomaly detection where applicable, alert routing, deduplication, and noise reduction • Integrate monitoring/alerting and events with BigPanda (correlation, enrichment, routing, incident workflows) • Create and maintain dashboards and operational visibility (Grafana or equivalent) • Develop and maintain runbooks, operational playbooks, and incident response procedures • Participate in on-call shifts: triage alerts, manage incidents, coordinate response, and lead communication during outages • Perform root-cause analysis, postmortems, and implement corrective/preventive actions • Improve service reliability via SLOs/SLIs, capacity planning, and automation to reduce toil • Support monitoring for core infrastructure and services on Windows and Linux, including HA components and clusters • Collaborate with DevOps/Engineering to instrument applications and standardize telemetry (metrics, logs, traces where applicable)

Job Requirements

  • Bachelor in Computer Science or related field
  • Experience in SRE / Operations / DevOps with production incident ownership
  • Hands-on experience with Prometheus and/or VictoriaMetrics (exporters, alert rules, recording rules, troubleshooting)
  • Experience integrating alerting/event pipelines with BigPanda (or similar event correlation tools)
  • Strong troubleshooting skills across Linux and Windows systems (networking, OS, services)
  • Ability to build reliable alerting with minimal noise (correlation, grouping, suppression, maintenance windows)
  • Experience with Git-based workflows for monitoring-as-code and configuration management
  • Nice to have
  • Grafana administration and dashboard design standards
  • Log management (ELK/EFK, Loki) and/or tracing (OpenTelemetry)
  • Automation skills (Python, PowerShell, Bash) and configuration tools (Ansible)
  • Messaging/cache/proxy operations: RabbitMQ, Redis, Nginx
  • Experience with Windows clustering or HA environments
  • Experience defining SLOs/SLIs and operational KPIs
  • Experience in managing VOIP components and protocols (SIP , FreeSwitch, OpenSIP, session border controllers)
  • Experience with load balancing components ( F5 LTM, F5 GTM)
  • Experience with Virtualization platforms such as VMWare or HyperV
  • Experience with administering AWS or Azure tenants

Benefits

  • We hire, promote, and compensate employees based on their ability to perform their job responsibilities, without regard to race, color, creed, religion, sex, gender, marital status, national origin, ancestry, age, citizenship, physical or mental disability, sexual orientation, or any other basis protected by applicable law (collectively referred to in our Code of Conduct as “Protected Classes”). We do not tolerate employment discrimination in the workplace, and we are committed to making reasonable accommodations for identified disabilities or other limitations as required by all applicable laws. We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.*

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Spotify logo

Senior Site Reliability Engineer

Spotify

Passionate music fans. Innovative tech pros. Perfect harmony. Join our band.

DevOps Engineer132 days ago
OtherRemoteTeam 5,001-10,000Since 2008H1B Sponsor

• Own fleet reliability. Lead the reliability, security, and scalability strategy for Portal’s SaaS infrastructure, including the runtime environments that power our platform and LLM-driven agent workflows. Define SLOs, drive capacity planning, and ensure our systems meet the demands of a rapidly growing product. • Architect for the agentic era. Design and evolve infrastructure on GCP and AWS using Terraform and infrastructure-from-code patterns. Shape how we structure environments for non-deterministic AI workloads — including sandboxing, resource isolation, cost governance, and security boundaries. • Drive operational excellence. Evolve our incident management, on-call, and postmortem practices. Leverage AI assistants to accelerate root cause analysis and build increasingly self-healing capabilities into our production systems. • Lead fullstack reliability. Operate across a modern web stack (TypeScript, React, Python). While not frontend-heavy, you’ll diagnose and resolve issues across the stack and drive reliability improvements end-to-end. • Mentor and multiply. Raise the reliability IQ of the broader engineering team. Establish SRE best practices, conduct production-readiness reviews, and mentor engineers on operational thinking. • Shape the roadmap. Partner with engineering and product leadership to evolve our infrastructure in step with generative AI features. Translate operational insights into strategic input on the product roadmap.

New York
$164.4K - $234.9K / year
Job Closed
F5 logo

DevOps Engineer III

F5

We secure every app.

DevOps Engineer132 days ago
Full TimeRemoteTeam 5,001-10,000H1B Sponsor

• Lead design, development, and optimization of CI/CD pipelines using industry-leading tools • Implement robust Infrastructure as Code (IaC) for scalable, secure on-prem and cloud environments • Drive automation of provisioning, deployment, testing, and monitoring leveraging scripting languages (Python, Bash, golang) • Develop and manage containerized applications and orchestrate workloads with Docker and Kubernetes • Utilize cloud platforms (AWS, Azure) to optimize infrastructure performance, resilience, and cost • Implement and support continuous monitoring, logging, and alerting utilizing tools such as Prometheus, Grafana, and OpenTelemetry (OTEL) • Apply configuration management solutions (Ansible or other - Chef, Puppet, or SaltStack) • Champion DevSecOps practices by integrating security into all automated pipelines and infrastructure • Lead troubleshooting and resolution of complex deployment, infrastructure, and system issues • Collaborate cross-functionally with software engineering, QA, product management, security, and IT operations throughout agile, cross-disciplinary teams • Establish and maintain clear, comprehensive documentation including procedures, runbooks, technical solutions to support team knowledge and operational excellence • Mentor, coach, and provide technical leadership to DevOps engineers and advocate for best practices across engineering teams

India
Job Closed
Full TimeRemoteTeam 10,001+Since 1978H1B No Sponsor

• Design and implement infrastructure automation using IaC tools. • Build and maintain CI/CD pipelines. • Manage and optimize GKE clusters and containerized workloads. • Develop internal developer platforms and self-service tools. • Implement monitoring, logging, and alerting for platform reliability. • Collaborate with engineering and operations teams to improve workflows. • Ensure platform security, compliance, and scalability via automation. • Document architecture, processes, and tooling.

Ukraine
Job Closed
Coderio logo

SRE, Cloud Engineer

Coderio

Accelerate Your Digital Transformation

DevOps Engineer132 days ago
Full TimeRemoteTeam 201-500Since 2017H1B No Sponsor

• Acompañar a los squads en la adopción de prácticas SRE • Prototipar soluciones técnicas innovadoras • Acelerar la adopción de estándares de calidad en toda la organización

Argentina
Job Closed