Take your entire business from reactive to proactive with the leading AI-Powered Complete FP&A Platform.
Site Reliability Developer
Location
Canada
Posted
3 days ago
Salary
$85K - $115K / year
Seniority
Mid Level
Job Description
Site Reliability Developer
Vena Solutions
• Support key ITIL processes, including Incident management, request management, problem management and change management. • Define and document runbooks and standard operating procedures. • Field operational requests from our Application Support team and other internal stakeholders • Triage and solve issues within defined SLA’s to ensure an excellent customer experience and to unblock other development and support teams • Maintain services once they are live by measuring and monitoring availability, latency and overall system health. • Identify and troubleshoot problems, investigate root causes, and champion fixes across the organization. • Work with infrastructure-as-Code (IaC) with a focus on continuous improvement. • Collaborate with cross-functional team members on features and implementation within an agile environment. • Report on SLAs and performance metrics as part of the Operations function. • Participate in on-call rotation.
Job Requirements
- Bachelor’s degree in computer science, Software engineering or equivalent experience
- 2+ years of experience in an IT Operational, DevOps, SRE, or Software Engineering role.
- Experience with cloud computing (AWS and Azure) services and a developing-level of knowledge with the management and setup of cloud infrastructure.
- You can write code - in any language. You have implemented your work in a production environment and can back it up with examples.
- Experience with tools and platforms such as: Ansible, Build/Release Pipelines, Docker, Github, Terraform etc.
- Developing-level of knowledge with distributed systems in the cloud using observability and telemetry for oversight of code deployments and service level objectives (SLOs).
- Developing experience with the operational aspects of software systems using telemetry, centralized logging, and alerting with tools such as: CloudWatch, Datadog, Prometheus, etc.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• proaktywne dostrzeganie wyzwań i ich adresowanie wspólnie z zespołami backendowymi i devops • tworzenie szablonów i standardów infrastrukturalnych dla najczęstszych potrzeb zespołów developerskich (np. nowe serwisy, projekty Cloud Run) • rozwój i utrzymanie stacku observability (Loki, Grafana, Prometheus) oraz optymalizacja wykorzystania zasobów i kosztów chmury (GCP) • utrzymanie i rozwój wewnętrznych narzędzi self-hosted (m.in. NocoDB, n8n, Outline) • wsparcie infrastrukturalne zespołów Data Engineering i Data Science (m.in. Airflow, pipeline'y forecastingowe czy serwery pod ML) • zarządzanie infrastrukturą sieciową w darkstore'ach (sieć, CCTV, drukarki paragonów, urządzenia handheld) we współpracy z IT Managerem • budowanie i utrzymanie infrastruktury on-prem (serwery pod ML i workloady wymagające dużych zasobów, runnery CI/CD dla projektów mobilnych) • utrzymanie i usprawnianie procesów CI/CD (GitLab) oraz developer experience • budowanie infrastruktury pod rozwiązania AI (autonomiczne agenty, boty Slack, agenty code review, remote coding agents) • rozwój praktyk SRE: alerting, procesy zgłaszania i obsługi incydentów (PagerDuty, Slack)
Senior DevOps Engineer
MKS2 TechnologiesAustin-based SDVOSB delivering application development, cybersecurity, instructional design and training to DOD and VA.
• Conduct DevOps and DevSecOps activities for an Azure integration platform • Create and maintain GitHub CI/CD pipelines • Conduct Azure DevOps activities • Conduct Azure system administration • Build platform automation • Create and maintain PowerShell scripting • Understanding and provisioning of Azure Infrastructure services and network • Create and maintain containers and infrastructure (IaC, Terraform) • Deploy solutions through CI/CD pipelines • Embed DevOps best practices and process/tech improvements across the team and platform • Leverage AI tools to accelerate activities • Collaborate across technical teams • Create and maintain detailed technical documentation • Support troubleshooting and resolution of Production issues • Participate in Agile ceremonies • Report on progress and status of development
Senior DevOps Engineer
FiservFounded in 1984, Fiserv is a global provider of ecommerce and information management systems for the financial services industry. In 2013, Fiserv acquired Open
• Design, implement, and support secure Azure IaaS and PaaS solutions across enterprise environments • Build and maintain CI/CD pipelines using Azure DevOps and GitLab to enable efficient, automated software delivery • Develop and manage Infrastructure as Code solutions using Terraform and Ansible for provisioning and configuration management • Implement and maintain Zero Trust security architectures, including identity and access management (IAM), data encryption, and security controls • Identify, assess, and remediate infrastructure and application security vulnerabilities • Create and maintain monitoring, alerting, and operational dashboards using Azure Monitor, Log Analytics, and Application Insights • Troubleshoot production issues, restore services, and maintain operational runbooks and knowledgebase documentation • Collaborate with cross-functional teams to support cloud modernization initiatives
Senior SRE, Managed Gateways
Kong Inc.The cloud connectivity company. Powering connections to build a reliable digital world.
• Lead, mentor, and inspire a high-performing team of Site Reliability Engineers dedicated to Kong's Managed Gateway offerings. • Architect and implement robust, scalable, and fault-tolerant cloud-native systems using technologies like Kubernetes, Golang, and major cloud providers. • Own the end-to-end operational lifecycle, from proactive monitoring and alerting to incident response and blameless post-mortems, ensuring continuous service improvement. • Drive a culture of developer delight by implementing automation, self-service tooling, and streamlined workflows for deploying and managing API gateways. • Define, track, and report on key SLOs and SLIs to ensure optimal performance and reliability of Managed Gateways. • Champion technical debt prevention and advocate for architectural best practices that enhance system resilience and reduce operational toil. • Collaborate cross-functionally with Product, engineering, and Customer Success to influence roadmap decisions and ensure operational readiness for new features. • Partner directly with enterprise customers — working alongside Product leadership, Professional Services, and Customer Success — to drive end-to-end onboarding and implementation of Cloud Gateways, and productize recurring implementation patterns into repeatable playbooks and platform capabilities. • Bring deep, cross-cloud breadth (AWS, GCP, Azure) to handle unique customer topologies and turn complex setups into successful, production-ready deployments.




