Our PK-12 solutions ignite the power of literacy through easy-to-implement programs rooted in the Science of Reading. ✳️
Senior DevOps Engineer
Location
Canada
Posted
12 days ago
Salary
0
Seniority
Senior
Job Description
Senior DevOps Engineer
Really Great Reading
• Ensure the stability, performance, and security of our AWS production environment through end of life • Monitor critical services, respond to incidents, and maintain high availability • Optimize AWS costs thoughtfully during the transition period • Handle day-to-day operations across GCP services — monitoring, security hardening, and observability • Maintain and improve CI/CD pipelines to support reliable, efficient deployments • Manage Infrastructure as Code (Terraform) with a focus on best practices, reviews, and continuous improvement • Support deployments and data migrations as needed • Promote and model best practices in IaC, GitOps, observability, and security • Document processes clearly and support teammates in understanding and using shared tools • Contribute to a culture of operational rigor, async collaboration, and continuous learning
Job Requirements
- 5+ years of experience in DevOps, SRE, or Cloud Engineering
- Strong expertise in AWS (EC2, RDS, VPC, IAM, S3, CloudWatch, and related services)
- Significant hands-on experience with GCP (Compute Engine, GKE, Cloud SQL, IAM, VPC, Cloud Logging)
- Proficiency with Terraform — modules, state management, and best practices
- Solid production experience with Docker and Kubernetes
- Strong understanding of cloud networking, security, and IAM
- Experience with at least one CI/CD platform (GitHub Actions, GitLab CI, CircleCI, or similar)
- Comfortable scripting in Bash, Python, or Go
- Excellent written communication and ability to thrive in a fully remote, async environment.
Benefits
- Remote: Quebec, Canada
- Competitive salary
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
DevOps Engineer
epensionZusammenhalt & Support: Wir sind ein Team, das sich trotz der Distanz gegenseitig unterstützt, Wissen aktiv teilt und sich bei der Lösungsfindung unter die Arme greift. Gelebte Wertschätzung: Ein hilfsbereites Umfeld und echte Wertschätzung füreinander sorgen dafür, dass du dich vom ersten Tag an einbringen kannst. Gemeinsam wachsen: Durch ein abgestimmtes Onboarding mit Mentoren sowie eine offene Feedbackkultur stellen wir sicher, dass wir als Team – trotz Remote-Fokus – eng zusammenwachsen.
Role Description Deine Stimme zählt! Verstärke unser DevOps-Team. Gestalte die zukünftige Infrastruktur aktiv mit und sei prägend im Aufbau der DevOps-Kultur. - Aufbau und Weiterentwicklung unserer Infrastruktur - Etablierung einer DevOps-Kultur - Crossfunktionale Zusammenarbeit und entwicklungsnahe Mitarbeit - Automatisierung und Standardisierung - Containerisierung und Migration - KI-Innovation und Prozessautomatisierung - Monitoring, Logging und Systemtransparenz - Mitgestaltung und Dokumentation - Strukturiertes Onboarding & individuelle Einarbeitung - Stabile Prozesse & moderne Technologien - Abwechslungsreiche Projekte mit technischem Tiefgang - Flexibles Arbeiten mit 100 % Homeoffice-Möglichkeit Qualifications - Mehrjährige Erfahrung im Bereich DevOps, Cloud-Integration oder Infrastruktur-Administration - Sichere Bewegung in Linux-Umgebungen (Ubuntu) mit fundierten Kenntnissen in Server-, Netzwerk- und Systemkonfiguration - Erfahrung mit Ansible, Docker und Proxmox; idealerweise Terraform oder andere IaC-Werkzeuge - Fähigkeit, Public- und Private-Cloud-Architekturen aufzubauen und weiterzuentwickeln - Nutzung von Monitoring- und Logging-Tools wie Grafana, Loki oder LibreNMS - Wohlfühlen in entwicklungsnahen Aufgaben und Verständnis von Kotlin-basiertem Code (inkl. Gradle-Builds) - Bereitschaft, DevOps-Prinzipien im Unternehmen zu verankern - Analytische, strukturierte und lösungsorientierte Arbeitsweise - Schätzung eines Arbeitsumfelds, in dem Ideen eingebracht und technische Entscheidungen mitgestaltet werden können Requirements - Du bringst Freude daran mit, Strukturen aufzubauen, nicht nur zu warten. - Du hast Lust, DevOps-Prinzipien im Unternehmen zu verankern. Benefits - Unbefristete Festanstellung - 30 Tage Urlaub - Flexible Arbeitszeiten - Zuschuss zur Altersvorsorge - Betriebliche Krankenzusatzversicherung - Dienstrad - Corporate Benefits - Essenszuschuss - Bezuschusstes Deutschlandticket - Workation-Möglichkeiten - Unabhängige Beratung in allen Lebenslagen - Weiterentwicklung nach deinen Vorstellungen Company Description Wir entwickeln digitale Lösungen rund um betriebliche Vorsorge. Komplexe Prozesse sollen so gestaltet sein, dass sie für alle Beteiligten – Arbeitgebende, Versicherungen und Beschäftigte – einfach, sicher und effizient ablaufen. Damit das gelingt, bauen wir unsere Infrastruktur und unsere Arbeitsweise konsequent weiter aus – hin zu einer modernen, Cloud-basierten und teamübergreifend automatisierten Plattform. Dafür suchen wir Menschen, die Lust haben, sich aktiv einzubringen, Verantwortung zu übernehmen und die DevOps-Kultur bei epension gemeinsam mit uns zu etablieren. Wir freuen uns darauf, dich kennenzulernen!
Systems Reliability Engineer
Bright Vision TechnologiesBright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.
Role Description We are seeking an experienced Site Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production. As an SRE you will live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with lower operational toil. The ideal candidate will combine deep systems knowledge with strong programming skills, a measurement-driven mindset, and the discipline to design, automate, and operate complex services so that reliability becomes a first-class engineering deliverable rather than a reactive concern. Key Responsibilities - Define, instrument, and continually refine service-level objectives (SLOs), service-level indicators (SLIs), and error budgets for critical services. - Lead incident response and resolution for production issues, acting as a calm and effective incident commander when needed. - Ensure high-quality post-incident reviews that drive lasting improvements. - Design and implement comprehensive monitoring, logging, and tracing strategies using tools like Prometheus, Grafana, OpenTelemetry, ELK/EFK, Datadog, or similar. - Build and maintain robust on-call processes, runbooks, and escalation paths. - Automate operational toil aggressively by writing production-grade tooling in Python, Go, Bash, or similar languages. - Architect and operate large-scale Kubernetes clusters and container-based workloads. - Design CI/CD pipelines that promote safe, frequent, and observable releases. - Lead capacity planning and performance engineering activities. - Partner closely with application development teams to embed reliability practices early in design. - Strengthen the platform’s resiliency through chaos engineering and fault injection. - Drive continuous improvement of security posture in collaboration with security teams. - Contribute to the technical roadmap for reliability tooling and observability platforms. - Mentor engineers across the organization on SRE practices. Qualifications - Bachelor’s degree in Computer Science, Engineering, or a related technical discipline. - Five or more years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems. - Strong programming skills in at least one of Python, Go, or Java. - Deep, hands-on experience operating Linux at scale. - Production experience operating Kubernetes and container-based workloads. - Strong working knowledge of observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or commercial equivalents. - Hands-on experience designing and operating CI/CD pipelines. - Solid understanding of distributed system design. - Demonstrated experience leading incident response and conducting effective post-incident reviews. - Excellent communication and documentation skills. Preferred Qualifications - Experience defining and operationalizing SLOs and error budgets in real production environments. - Exposure to chaos engineering practices and tools such as Chaos Monkey, Gremlin, or Litmus. - Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP). - Background in capacity planning, performance engineering, or large-scale load testing. - Familiarity with service mesh technologies such as Istio, Linkerd, or Consul. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 650-6699. Learn more about Bright Vision Technologies at www.bvteck.com . Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
Corporate Reliability Engineer
ArclinOur vital technologies are everywhere, reinforcing products the world can’t live without.
• Develop and implement strategies to enhance the reliability and performance of manufacturing processes and systems. • Analyze data, identify root causes of issues, and design solutions to improve equipment reliability. • Conduct reliability assessments and failure analyses to identify weaknesses in systems and processes. • Develop, implement, and govern predictive maintenance and condition-monitoring strategies for equipment. • Perform detailed analyses of equipment failures and process breakdowns using FMEA and RCA. • Lead continuous improvement initiatives in equipment reliability and process efficiency. • Work closely with cross-functional teams to address reliability issues and implement solutions.
• Implement security practices in software development (DevSecOps); • Automate security checks in CI/CD pipelines; • Monitor and mitigate vulnerabilities in applications and infrastructure; • Perform static and dynamic code analysis to identify security flaws; • Implement and configure security tools for applications and infrastructure; • Collaborate with development, infrastructure, and security teams; • Define and maintain security and compliance policies; • Promote security training and awareness for technology teams; • Respond to security incidents and propose corrective measures; • Participate in agile ceremonies (daily meetings, planning, review, and retrospective).


