Manage your maintenance and operations without the paper stacks.
Site Reliability Engineer
Location
Canada
Posted
1 day ago
Salary
0
Seniority
Senior
Job Description
Site Reliability Engineer
MaintainX
• Assess service maturity and provide insights to development teams • Partner with development teams to implement observability best practices • Enable development teams to become autonomous with their service deployment, support, and infrastructure • Mentor developers on reliability practices, focusing on making them self-sufficient • Act as the bridge, ear and eyes of the Platform Division teams to drive tooling and practice adoption across development teams
Job Requirements
- Deep understanding of observability practices in a distributed system environment and how it influences system design and team behaviour
- Practical experience with SRE concepts (SLOs, error budgets, incident management)
- 3–5+ years in software development, SRE, DevOps, or production development roles with experience operating production systems
- Proficient in cloud-native platforms and infrastructure-as-code concepts and tools
- Working knowledge of at least one programming language (TypeScript/Node.js is a plus)
- Excellent communication and collaboration abilities across technical and non-technical teams
- Ability to translate complex reliability concepts into actionable guidance
- You enjoy enabling teams to succeed independently and measuring success by reduced dependency on you.
Benefits
- Competitive salary and meaningful equity opportunities.
- Healthcare, dental, and vision coverage.
- 401(k) / RRSP enrollment program.
- Take what you need PTO.
- A Work Culture where:
- You’ll work alongside folks across the globe that reflect the MaintainX values, Smart Humble Optimist.
- We believe in meritocracy, where ideas and effort are publicly celebrated.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Design, develop, test, deploy, and maintain automation solutions using Microsoft Power Automate (cloud and desktop flows). • Build workflows that improve efficiency, reduce manual effort, and increase process consistency. • Develop scalable, supportable solutions using approved Power Platform tools and patterns. • Contribute to enhancements and continuous improvement of existing automations. • Partner with stakeholders to understand processes, pain points, and business goals. • Document current-state and future-state processes, translating requirements into technical designs. • Identify opportunities for automation, simplification, and standardization. • Recommend practical, user-friendly solutions aligned with operational needs. • Monitor automation performance and troubleshoot production issues promptly. • Perform root cause analysis, defect resolution, and ongoing maintenance. • Develop solutions in line with Power Platform governance, security, and compliance standards. • Create and maintain technical documentation, support guides, test plans, and release notes. • Collaborate with analysts, developers, business stakeholders, and technology partners globally.
Site Reliability Engineer
The Voleon GroupApplying statistical machine learning to investment management.
• Improve fault-tolerance and maintainability of code in proprietary data pipelines and trading systems • Diagnose and fix bugs in code • Lead complex deployments • Automate manual workflows • Track and prioritize outstanding production-related issues • Share an on-call rotation responding to incidents to ensure the continuous operation of production-critical systems
Senior Site Reliability Engineer
The Voleon GroupApplying statistical machine learning to investment management.
• Be a first responder in the event of cluster outages or issues. Triage and resolve urgent issues as they arise • Ensure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliability • Diagnose systemic/recurring patterns of problems, and engineer precision solutions to them in collaboration with engineering teams • Develop robust metrics and observability for cluster health and use those metrics to inform your work. Build out custom observability mechanisms when off-the-shelf ones won't do • Help software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policies • Assist in forecasting cluster growth, and help select appropriate scale-up strategies. Help optimize operations across dimensions of cost and usability
DevOps Engineer
CareerswiftJob Searching Shouldn't Feel Like a Full-Time Job. AI-Powered Career Acceleration That Works
• Designing, automating and maintaining CI/CD pipelines • Cloud infrastructure (AWS/GCP) • Monitoring systems • Improve deployment reliability, scalability and security across platform



