Chainlink Labs logo
Chainlink Labs

Chainlink is the industry-standard oracle platform bringing the capital markets onchain and powering the majority of decentralized finance (DeFi). The Chainlink stack provides the essential data, interoperability, compliance, and privacy standards needed to power advanced blockchain use cases for institutional tokenized assets, lending, payments, stablecoins, and more. Since inventing decentralized oracle networks, Chainlink has enabled tens of trillions in transaction value and now secures the vast majority of DeFi. Many of the world’s largest financial services institutions have also adopted Chainlink’s standards and infrastructure. Chainlink leverages a novel fee model where offchain and onchain revenue from enterprise adoption is converted to LINK tokens and stored in a strategic Chainlink Reserve.

Senior Site Reliability Engineer, DevEx

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 201-500Since 2017H1B No SponsorCompany SiteLinkedIn

Location

Canada

Posted

6 days ago

Salary

$90K - $125K / year

Seniority

Senior

Job Description

Senior Site Reliability Engineer, DevEx

Chainlink Labs

• design and build the infrastructure primitives that define how our CI/CD platform, build systems, and developer environments scale across the entire engineering org. • help build and operate the Kubernetes-based control plane behind our CI/CD platform, including: - GitHub Actions self-hosted runner infrastructure (autoscaling, isolation, cost/perf tuning) - GitHub Apps and GitHub-as-code (permissions, webhooks, org-wide automation) - Secure network access for CI/CD and remote dev environments (Tailscale) - GitOps-driven deployment of platform services (Flux) - Ephemeral/on-demand developer environments and build systems • develop the core infrastructure components — including Kubernetes Operators and scaling automation — that product teams adopt directly, reducing bespoke per-team CI/CD and environment tooling. • building the systems that define how engineering teams build, test, and deploy, shaping the reliability and scalability of the developer experience org-wide.

Job Requirements

  • 6–9+ years in SRE / Platform / Infrastructure Engineering
  • Proven experience scaling Kubernetes in high-throughput production environments
  • Deep Kubernetes expertise beyond operating clusters — internals, scheduler behavior, custom resources, and cluster-scale failure diagnosis
  • Experience building platform infrastructure, control planes, or Kubernetes Operators (not just consuming them)
  • Strong distributed systems and production reliability experience
  • Terraform/GitOps ownership — designing and owning automation, not just running playbooks
  • GitOps workflows (Flux / ArgoCD) experience
  • Hands-on experience with CI/CD platforms at scale: GitHub Actions (self-hosted runners, workflows-as-code), GitHub Apps, and build systems
  • AWS/cloud infrastructure production experience
  • Proficiency in Go (strongly preferred) or another systems language
  • Track record of building infrastructure primitives rather than primarily performing support/operations — automation-first mindset.

Benefits

  • comprehensive benefits
  • long-term incentives

Related Categories

Related Job Pages

More DevOps Engineer Jobs

AITASTIC AG logo

DevOps Engineer

AITASTIC AG

Next Level Consumer and Communication Insights

DevOps Engineer6 days ago
Full TimeRemoteTeam 51-200H1B No Sponsor

• Operate and deploy applications within an SOA architecture (GCP, Kubernetes, Terraform) • Ensure high availability of databases (Elasticsearch, Redis, and vector databases) • Optimize logging and monitoring using Grafana and Prometheus • Contribute to the design and implementation of security concepts

Germany

Role Description We’re looking for a DevOps Engineer to join our fully-remote team and help architect, deploy, and operate mission-critical cloud platforms that support everything from financial services to healthcare and government systems. This is a hands-on role where you’ll be working at the intersection of automation, Kubernetes, GitOps, and sovereign cloud infrastructure — helping organisations transition from legacy systems to resilient, containerized environments built for performance, security, and data sovereignty. If you're passionate about CI/CD, infrastructure-as-code, and building scalable deployment workflows that actually make developers’ lives easier — you’ll feel right at home here. What You’ll Be Working On - Designing and maintaining automated deployment pipelines and cloud-native platforms running on fully sovereign infrastructure. - Supporting clients through: - Infrastructure migrations - Platform modernisation - High-availability cluster deployments - Secure containerised workloads - Observability and incident response - Working closely with development teams to streamline delivery pipelines and implement scalable deployment strategies across multi-tenant OpenShift platforms. Your Impact - Design, implement, and maintain CI/CD pipelines to streamline build and release processes. - Install, configure, and manage OpenShift clusters across staging and production environments. - Deploy and manage storage platforms including: - OpenShift Data Foundation (ODF) - Ceph - Optimise Kubernetes, Docker, and OpenShift environments for performance, availability, and security. - Automate infrastructure provisioning using: - Terraform - Ansible - FluxCD - Develop Helm charts and Kustomize configurations within a GitOps-driven deployment model. - Monitor system performance and troubleshoot production incidents. - Implement monitoring and logging solutions using: - Prometheus - ELK Stack - Deploy and manage security integrations: - TLS Certificates - Authentication via Keycloak - ISO-compliant security frameworks - Manage containerised database environments: - PostgreSQL - MySQL - Oversee network, storage, and data centre integrations. - Automate recurring infrastructure tasks and optimise platform workflows. - Support customers directly across planning, deployment, and operational phases. - Document systems and contribute to internal knowledge-sharing initiatives. Qualifications - Solid experience with modern CI/CD practices. - Hands-on expertise in: - Kubernetes - Docker - Helm - Kustomize - OpenShift - Experience with GitOps platforms: - FluxCD - ArgoCD - Familiarity with storage platforms such as: - OpenShift Data Foundation - Ceph - Infrastructure-as-Code experience with: - Terraform - Ansible - Scripting in: - Bash - Go or Python - Experience managing relational databases: - PostgreSQL - MySQL - Experience working with cloud and on-premise environments. - Monitoring and logging with: - Prometheus - ELK Stack - Strong familiarity with: - Git - Jira - Confluence - Knowledge of networking, storage, and data centre operations. - Experience with security standards, TLS, and ISO compliance. - Excellent communication skills and the ability to work independently in a remote environment. Benefits - Work on enterprise-scale infrastructure running on cutting-edge hardware — including NVIDIA GPU clusters supporting AI training, inference, and HPC workloads. - Ensure every deployment meets the highest standards of Swiss data protection and sovereignty. - Empowered to propose, build, and ship solutions that balance innovation with real-world compliance and operational resilience.

Worldwide
Glia logo

Senior Software Engineer – SRE, Observability Tooling

Glia

The #1 Platform for Intelligent Banking Interactions

DevOps Engineer6 days ago
Full TimeRemoteTeam 201-500Since 2012H1B No Sponsor

• Focus on building SRE and observability tooling — the platform, automation, and standards other teams use to keep their services healthy. • Developing standards, infrastructure and automation for dashboards, alerts, and monitors as code. • Partnering with development teams to establish production readiness and operational readiness. • Building the tooling and templates teams use to define, measure, and report on Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for their services. • Developing tooling to automate observability and operational workflows, eliminating manual toil for engineering teams. • Building and improving the incident response tooling and workflows that help teams resolve outages faster and learn from them.

Estonia
Glia logo

Senior Software Engineer, SRE / Observability Tooling

Glia

The #1 Platform for Intelligent Banking Interactions

DevOps Engineer6 days ago
Full TimeRemoteTeam 201-500Since 2012H1B No Sponsor

• Build SRE and observability tooling • Develop standards, infrastructure, and automation for dashboards and alerts • Partner with development teams for production and operational readiness • Build tooling and templates for defining and reporting on Service Level Objectives (SLOs) • Automate observability and operational workflows • Improve incident response tooling and workflows

Poland