Strategic open source infrastructure for containers and virtual machines.
Senior Site Reliability Engineer, Kubernetes
Location
Canada
Posted
1 day ago
Salary
0
Seniority
Senior
Job Description
Senior Site Reliability Engineer, Kubernetes
Mirantis
• Work with geographically distributed international teams on technical challenges and process improvements. • Develop, implement, maintain, and troubleshoot cloud and AI infrastructure solutions based on open source software. • Collaborate with stakeholders to gather and refine technical requirements. • Optimize system performance, reliability, and scalability. • Troubleshoot, debug, and resolve complex technical issues. • Participate in code reviews to maintain high quality standards. • Stay up to date with industry trends and best practices in cloud operations and development. • Design and implement AI-driven automation across the DevOps lifecycle, including code development and maintenance. • Facilitate knowledge transfer to customers during the delivery phases.
Job Requirements
- 5+ years of professional experience in DevOps, with a strong focus on Cloud, infrastructure technologies and Kubernetes
- Experience with high-performance data center processing, networking, and storage
- Exposure to Golang and working knowledge of other programming languages (Python, JavaScript).
- Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines.
- Exceptional problem-solving and debugging skills across networking and storage (hardware and software), Linux, and Kubernetes, with attention to performance optimization and security.
- Demonstrated ability to lead technical tasks and collaborate effectively with diverse teams.
- Comfortable making independent judgment calls when working directly with customers, often with limited day-to-day oversight.
- Excellent written and spoken English.
- Excellent customer-facing communication skills.
- A commitment to innovation, continuous learning, and delivering high-quality results.
- Ability to travel up to 25% if needed, including internationally.
Benefits
- Work with an established Silicon Valley leader in the cloud infrastructure industry;
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
- Be a part of cutting-edge, open-source innovation;
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
- Professional development and training;
- Attend conferences and working groups;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Ensuring the reliability, availability, and performance of production infrastructure and platform services • Operating and scaling Kubernetes platforms, including governance and support for multi-tenant workloads • Managing GitOps-based deployment workflows using ArgoCD and Helm • Driving infrastructure provisioning and change management through Terraform/Terragrunt • Building and supporting CI/CD automation and deployment workflows using GitHub Actions • Leading incident response efforts, root cause analysis, and post-incident improvement initiatives • Reducing operational toil through scripting, tooling, and process automation • Advancing observability practices across logs, metrics, traces, dashboards, and alerting • Supporting secure secrets integration, IAM-aware operations, and platform guardrails • Partnering closely with application, security, and platform teams to improve reliability and delivery outcomes
Software/Site Reliability Engineer – FedRAMP
TenableCloud Security | Operational Technology | Identity Security | and more
• Responsible for taking the code and functionality of Tenable cloud products and ensuring they’re reliable and highly available in cloud environments • Responsible for responding to support escalations which involve troubleshooting complex technical problems and resolving data/configuration issues within defined service level objectives • Responsible for developing software, tools, and scripts to automate deployment, management, and monitoring of production systems in all environments • Collaborate with peers on complex projects • Collaboration with cloud engineers in understanding new cloud technologies, assessing impact to security services operations, and proposing solutions to existing business problems • Collaboration in the software development lifecycle to develop detailed enhancement/bug definitions, write functional requirements, translate the requirements into solution designs, and navigate the functional requirements through to Production deployments • Proactively look for ways to create efficiencies within operations as it pertains to the tools and technology used by Tenable to support their customer base • Manage, participate in, or directly work on any additional projects, assignments, or initiatives assigned by management • Create/maintain documentation for operational procedures • Contribute to standardization efforts across disciplines and services in conjunction with embedded SREs throughout the organization • Document and perform system upgrades, application updates, and define monitoring requirements based on customer or organization needs • Participate in an on-call rotation and support 24x7 availability of production application systems • Remediate infrastructure and/or container security vulnerabilities
DevOps Developer II
CoRelationBoutique strategy and communications agency focusing on corporate communications, public affairs and change.
• Support automation and lifecycle management of Corelation’s patching, upgrade, and system-maintenance tools. • Design, build, and maintain scripts and automated processes used by internal teams—and by clients—to upgrade, patch, and maintain RHEL-based environments. • Work collaboratively across Development, Hardware Services, and Client Support teams to build reliable and efficient processes for patching, upgrades, and system transitions. • Evaluate vendor updates and develop automation for internal and client use. • Maintain Corelation’s internal test systems and contribute to new projects requiring custom scripts or automation solutions.
Senior Site Reliability Engineer, SRE
MirantisStrategic open source infrastructure for containers and virtual machines.
• Work with geographically distributed international teams on technical challenges and process improvements. • Develop, implement, maintain, and troubleshoot cloud and AI infrastructure solutions based on open source software. • Collaborate with stakeholders to gather and refine technical requirements. • Optimize system performance, reliability, and scalability. • Troubleshoot, debug, and resolve complex technical issues. • Participate in code reviews to maintain high quality standards. • Stay up to date with industry trends and best practices in cloud operations and development. • Design and implement AI-driven automation across the DevOps lifecycle, including code development and maintenance. • Facilitate knowledge transfer to customers during the delivery phases.



