Knotch logo
Knotch

Marketing technology built to drive business outcomes for Content, Demand Gen, & Growth Marketers.

DevOps Engineer

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 51-200Since 2013H1B SponsorCompany SiteLinkedIn

Location

Canada

Posted

91 days ago

Salary

$90K - $120K / year

Seniority

Senior

Job Description

DevOps Engineer

Knotch

• Design, build, and maintain scalable, secure, and highly available infrastructure across pre-production and production environments • Develop and manage CI/CD pipelines to enable fast, reliable, and repeatable deployments across multiple environments • Own infrastructure as code (IaC) practices using tools like Terraform to ensure consistency and reproducibility • Manage environment lifecycle (development, staging, production), including promotion workflows and configuration management • Partner closely with Engineering, Data, and AI teams to support system performance, reliability, and scalability • Implement and maintain monitoring, logging, and alerting systems to ensure high visibility into system health and performance • Optimize infrastructure for cost, performance, and reliability, especially for compute- and data-intensive AI workloads • Support Kubernetes-based deployments and container orchestration for distributed systems • Contribute to security best practices across infrastructure, including IAM, networking, and application-level protections • Create dashboards and reporting systems to provide visibility into system performance, uptime, and operational metrics • Document architecture, operational processes, and infrastructure decisions to support knowledge sharing and onboarding • Act as a DevOps/SRE partner across teams, helping troubleshoot issues and improve system reliability

Job Requirements

  • 5+ years of experience in DevOps, Site Reliability Engineering, or Infrastructure Engineering roles within SaaS, PaaS, or cloud-native environments
  • Prior experience in growth-stage and/or startup environment scaling from $10M to $20M+ ARR with a lean team
  • Strong experience with Google Cloud Provider (GCP), including IAM, networking, and data services
  • Hands-on experience with Infrastructure as Code tools such as Terraform
  • Experience building and maintaining CI/CD pipelines (GitHub Actions, ArgoCD, or similar)
  • Solid experience with Kubernetes, Docker, and containerized environments
  • Familiarity with deployment tools such as Helm
  • Experience with monitoring and observability tools like Prometheus and Grafana
  • Strong understanding of system reliability, scalability, and performance optimization
  • Ability to work across multiple systems and priorities in a dynamic environment
  • Strong documentation and communication skills, with attention to clarity and detail
  • Supplementary experience supporting AI/ML or data-intensive workloads in production environments (Nice-to-Have)
  • Familiarity with workflow orchestration or data pipeline tools (Nice-to-Have)
  • Experience with cost optimization strategies for cloud infrastructure (Nice-to-Have)
  • Exposure to security frameworks and compliance best practices (Nice-to-Have)
  • Experience working with distributed or globally deployed systems (Nice-to-Have)

Benefits

  • Comprehensive medical, dental, and vision insurance eligibility
  • 401(k) plan
  • Unlimited PTO
  • 10+ company-paid holidays
  • A daily company-wide break, and more!

Related Categories

Related Job Pages

More DevOps Engineer Jobs

ultima milla logo

DevOps Engineer

ultima milla

Logistic Management System for E-commerce & Retail in Mexico. Raised +$7M USD from Y Combinator, FJLabs, & more.

DevOps Engineer91 days ago
Full TimeRemoteTeam 51-200H1B No Sponsor

• Desarrollar características y mejoras a los servicios y productos de forma segura y probada, de acuerdo con los lineamientos de la empresa. • Redactar documentación técnica. • Resolver problemas técnicos de alto alcance y complejidad. • Garantizar las mejores prácticas para servicios y productos de alta escala, con nuestro estilo de código y la mantenibilidad necesaria. • Solucionar problemas de rendimiento y optimización, particularmente a gran escala, así como demostrar la capacidad para diagnosticarlos y prevenirlos. • Asesorar y guiar a los miembros del equipo para ayudarlos a desarrollar sus habilidades técnicas, eliminando los obstáculos a su autonomía y agilidad. • Agregar valor al equipo y comunicar efectivamente tus ideas a tu líder de área. • Disponibilidad para resolver problemas críticos según sea necesario. • Colaboración para atender las actividades de soporte requeridas por nuestras áreas internas.

Argentina

Role Description This is a remote position. This role is ideal for engineers with 3-5 years of experience and a strong background in building secure, scalable platforms. We are looking for hands-on DevOps and Backend Engineers with real-world experience in: - Application/feature development - System design - Testing practices such as TDD - Full-stack development - Handling production incidents - Distributed systems - Modern infrastructure challenges What You’ll Do as a Software Craftsperson: - Design and document real-world DevOps and backend scenarios based on production incidents such as outages, scaling challenges, and secure deployments - Translate real engineering experiences into benchmark tasks that contribute to training next-generation AI systems - Contribute to building secure, scalable, Kubernetes-native architectures across modern infrastructure environments - Work across critical engineering domains including CI/CD pipelines, observability, identity & access management, infrastructure-as-code, and backend services - Collaborate with internal teams to design and simulate realistic engineering workflows and system behaviors - Apply practical engineering judgment to model distributed systems challenges and improve system resilience and reliability Qualifications - 3-5 years of experience in DevOps and Backend Engineering with a strong foundation in building secure, scalable systems - Strong hands-on expertise in DevOps and backend technologies (Node.js/Java/Python) including: - Kubernetes - Terraform - CI/CD pipelines - Tools such as k9s, k3s (GitLab CI preferred) - Backend technologies such as Python, Node.js or Java - Experience with Docker, gRPC, and Kubernetes-native services - Demonstrated experience working with secure, offline or air-gapped deployments (highly preferred) - Familiarity with distributed systems and backend architecture, with exposure to ML or distributed pipelines being a plus - Hands-on experience across multiple core functional areas, with exposure to at least five of the following: - Identity & Access Management - Observability (Prometheus + Grafana) - CI/CD Pipelines - Keycloak - GitLab CI - Terraform OSS - Kubernetes ecosystem tools - Strong problem-solving ability with real-world experience in handling production systems, incidents, and infrastructure challenges - Ability to work across multiple layers of the stack, from infrastructure to backend services, while ensuring scalability, reliability, and security Benefits Life at Incubyte: - We are a remote-first company with structured flexibility. Teams commit to shared rhythms during core hours, ensuring smooth collaboration while maintaining autonomy. - Twice a year, we come together in person for a co-working sprint and once a year for a retreat - with all travel expenses covered. - Our environment is built for crafters: experimenting with real-world systems, solving complex infrastructure challenges, and contributing to cutting-edge AI initiatives. - We are all lifelong learners, and our work is our passion. Perks: - Dedicated learning & development budget - Sponsorship for conference talks - Comprehensive medical & term insurance - Employee-friendly leave policies - Home Office fund - Medical Insurance

Worldwide
Job Closed
Full TimeRemoteTeam 501-1,000

Role Description In this role, you will be responsible for designing, operating, and evolving Hostinger’s core infrastructure to ensure scalability, reliability, and performance across our global platform. You will focus on: - Provisioning and maintaining systems - Improving automation - Enhancing observability - Ensuring services run efficiently at scale You will work on: - Optimizing distributed systems - Improving performance and resilience - Maintaining high availability across globally deployed services A key part of your role will involve building: - Reliable automation - Refining monitoring and alerting - Troubleshooting complex infrastructure issues to maintain maximum uptime As part of your responsibilities, you will also contribute to the operation and improvement of critical networking components, including DNS, ensuring consistent and low-latency service delivery worldwide. Working closely with product and infrastructure teams, you will help: - Evolve system architecture - Strengthen security - Support new features that enhance customer experience Your work will ensure fast, stable, and secure service delivery for millions of users globally. Qualifications - Advanced experience with Linux operating systems (Debian and Red Hat) - Solid hands-on experience with PDNS, MySQL clustering and replication, LMDB backends, DNSSEC - Proven ability to design, implement, and optimize NetBox, Ansible Tower, Terraform solutions - Skilled in monitoring system performance and conducting capacity planning - Proficient with automation tools such as Ansible - Strong problem-solving skills in diagnosing and resolving issues - Familiar with backup strategies and disaster recovery planning - Experience with monitoring tools to track system health and performance metrics - Dedicated to meeting Service Level Objectives (SLOs) - Always seeking opportunities to innovate and enhance existing systems - Strong communication skills in Lithuanian and English Requirements - Deploy and configure new systems, servers, and related components - Continuously manage and scale infrastructure to support increasing traffic - Build efficient, resilient systems including distribution models and redundancy strategies - Track system performance and ensure smooth service across all regions - Use tools like Ansible and Terraform to automate provisioning and deployment processes - Quickly diagnose and resolve infrastructure-related incidents - Continuously measure and fine-tune response times and system efficiency - Work closely with cross-functional teams to deliver infrastructure solutions - Explore new technologies and implement improvements Benefits - Limitless learning opportunities with access to Reforge, Couch Hub, global conferences, libraries, and mentoring - Work from modern offices or anywhere in the world with a flexible schedule - Health insurance from Day 1, gym memberships, recharge leave, and regular health checks - Company events like Summerfest & Winterfest, team-buildings, and milestone gifts Compensation Gross salary from 5000 EUR/month. Specific compensation is offered based on work experience, competence, and compliance with other job requirements.

Worldwide
€5K / month
Job Closed
Full TimeRemoteTeam 501-1,000

Role Description We are looking for a skilled DevOps engineer to take ownership of the DevOps domain. You will be the go-to person for infrastructure, CI/CD, and operations enabling the team to build, deploy, and operate reliably at scale. Your day-to-day responsibilities include: - Own DevOps responsibilities within a multi-product DevOps team - Manage and evolve Terraform-managed infrastructure on GCP and AWS - Operate and enhance applications deployed to Kubernetes - Collaborate with developers and the Core CloudOps team to implement DevOps best practices and ensure production-readiness - Ensure high availability, observability, and cost-efficiency across infrastructure components - Continuously improve and monitor CI/CD pipelines to support fast and safe delivery - Handle incidents, perform root cause analysis, and lead post-mortems related to the product’s infrastructure - Drive automation and self-service capabilities to reduce manual toil and increase team autonomy Qualifications - Experience: At least 5 years managing Linux systems - Technical Skills: Proficiency with Ansible, Terraform, Kubernetes, Docker, Git, SQL and CI/CD tools - Programming: Python, Go or similar languages - Problem-Solving: Strong analytical and proactive approach to complex technical challenges - Communication: Collaborative mindset and a high sense of ownership and accountability - Cloud Operations: Experience in highly available and scalable cloud environments in GCP or AWS - Kubernetes Workflows: Familiarity with Helm, Kustomize, or GitOps Requirements - Nice To Have: Knowledge of cost optimization techniques in cloud platforms - Experience in leading company-wide initiatives Benefits - 360 Growth: Limitless learning opportunities, access to platforms like Reforge and Coach Hub, global conferences, physical and digital libraries, feedback culture, and mentoring through TesoXchange. - Freedom & Responsibility: Work on your terms from modern offices or anywhere in the world with flexibility in managing your schedule. - Wellness Simplified: Health insurance from Day 1, gym memberships, recharge leave, Headspace subscriptions, and regular health checks. - Work Hard – Play Hard: Celebrate achievements with company events, access to the Žalgiris Arena VIP Lounge, and milestone gifts for weddings, new parenthood, and graduations. Compensation Gross salary from 5000 EUR/month. Specific compensation is offered based on work experience, competence, and compliance with other job requirements. We’re open to adjusting compensation based on the impact and value you bring.

Worldwide
€5K / month
Job Closed