Job Closed
This listing is no longer active.
The leading provider of enterprise open source solutions.
Senior Site Reliability Engineer – Golang, OpenShift, AWS, Linux
Location
Australia
Posted
202 days ago
Salary
0
Seniority
Senior
Job Description
Senior Site Reliability Engineer – Golang, OpenShift, AWS, Linux
Red Hat
• Develop, scale, and operate OpenShift managed cloud services • Enable customer self-service and improve monitoring systems • Eliminate work through automation • Participate in a regular on-call schedule, including occasional paid weekends and holidays • Resolve customer issues escalated from the Red Hat Global Support team • Work within a small agile team to develop and improve SRE software
Job Requirements
- A bachelor's degree in Computer Science or a related technical field is required
- Experience programming in at least one of these languages: Python, Golang, Java, C, C++ or another object-oriented language
- Experience working with public clouds such as AWS, GCP, or Azure
- Ability to collaboratively troubleshoot and solve problems in a team setting
- Experience troubleshooting an as-a-service offering (SaaS, PaaS, etc.)
- Experience working with complex distributed systems
- Direct experience with Kubernetes or OpenShift is a plus
- Demonstrated ability to debug, optimize code and automate routine tasks
- 5+ years of experience managing Linux servers running Red Hat Enterprise Linux (RHEL), CentOS, or Fedora hosted at a cloud provider
- 3+ years of experience with enterprise systems monitoring; knowledge of Prometheus is a plus
- 3+ years of experience with enterprise configuration management software like Ansible by Red Hat, Puppet, or Chef
- 2+ years of experience with at least one object-oriented language; Golang, Java, or Python are preferred
- 2+ years of experience delivering a hosted service
- Solid understanding of standard TCP/IP networking and common protocols like DNS and HTTP
- Solid communications skills and experience working directly with and presenting to customers
- 1+ year(s) of experience with Kubernetes is a plus
- 1+ year(s) of experience with docker-based containers is a plus
Benefits
- Health insurance
- Paid time off
- Flexible working hours
- Professional development opportunities
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Senior DevOps Engineer
eSimplicityAn engineering firm that delivers high-quality Healthcare IT, Cybersecurity, and Telecommunication solutions.
• Design, build, and maintain secure CI/CD pipelines using GitHub Actions to deliver applications and infrastructure • Embed security controls, tools (SAST, DAST, SCA), and processes throughout the software development lifecycle • Manage and secure cloud infrastructure using Infrastructure as Code (IaC) with Terraform and Terragrunt • Implement and manage security for containerized applications using Docker • Collaborate with development teams (Java, Python, Django) to identify and remediate security vulnerabilities in code and dependencies • Automate security monitoring, logging, and incident response procedures within the AWS cloud environment • Ensure systems and applications meet federal compliance standards (e.g., FISMA, NIST) and CMS-specific security requirements • Support the security of data platforms and services, including Databricks and Redshift • Work with cross-functional teams to foster a culture of security awareness and best practices
• Design, implement, and manage scalable cloud infrastructure on AWS. • Develop and maintain CI/CD pipelines using Bitbucket Pipelines. • Automate deployment processes. • Utilize Terraform for provisioning and managing cloud resources. • Ensure infrastructure is versioned, reproducible, and consistent across environments. • Drive the transition from traditional DevOps practices to platform development. • Develop and maintain internal platforms and tools to support development teams. • Implement monitoring solutions to ensure system reliability and performance. • Analyze and optimize system performance, identifying and resolving bottlenecks. • Collaborate with multiple development, security, and operations teams. • Implement security best practices and ensure compliance with industry standards. • Conduct regular security audits and vulnerability assessments. • Maintain comprehensive documentation of infrastructure, processes, and tools. • Share knowledge and mentor junior team members to foster a culture of continuous learning.
Build & Release Engineer
LeagueA platform technology company powering next-generation healthcare consumer experiences
• Design, maintain, and evolve automated build, packaging, and release workflows across web, backend, and mobile platforms • Own the end-to-end release lifecycle, including versioning, promotion, rollback, and hotfix strategies • Ensure high reliability and uptime of CI/CD systems through monitoring, alerting, and incident response • Support engineering teams in landing critical releases and resolving build or deployment issues • Build and optimize CI/CD pipelines using modern tooling (e.g., GitHub Actions, CircleCI, self-hosted runners) • Improve build performance, cost efficiency, and execution times across pipelines • Manage and evolve containerized build environments (Docker) and Kubernetes-based infrastructure • Implement infrastructure-as-code practices (e.g., Terraform) for reproducible build environments • Design and implement AI-assisted workflows to automate release validation, failure triage, and incident analysis • Optimize pipeline performance and identify bottlenecks • Generate and maintain build scripts, configs, and test scaffolding • Build AI-powered tooling (agents, bots, or copilots) for release orchestration and CI/CD diagnostics • Integrate AI into developer workflows to reduce manual effort and improve productivity • Continuously evaluate and adopt new AI tools and patterns to evolve League’s delivery capabilities • Build and maintain internal tools, dashboards, and CLI utilities to improve developer workflows • Enhance observability of build and release systems (logs, metrics, alerts) • Create and maintain documentation and self-service tooling for engineering teams • Partner with developers to understand friction points and improve release experience at scale • Implement and maintain secure build and release practices, including certificate and secrets management • Partner with Security and Platform teams to ensure safe and compliant releases • Improve reliability through incident management, postmortems, and continuous improvement loops
Site Reliability Engineer
DittoReal-time database for mobile, web, IoT, and server apps that can magically sync data with or even without the internet.
• Develop and maintain observability solutions using platforms like Datadog, Prometheus and Grafana • Take a leading role in incident management, including coordinating response efforts, troubleshooting issues, and identifying follow-up actions • Partner with product engineering teams to architect reliable systems, recover from incidents, and learn from mistakes • Work with teams to implement and maintain SLOs, monitoring, and alerting strategies that ensure reliability at scale • Design and implement automation and support tooling to improve system resilience, maintain operational safety and reduce operational overhead • Lead the development and maintenance of runbooks, alert definitions, and incident response procedures • Participate in on-call rotations to provide 24/7 support for critical production systems




