Caris Life Sciences logo
Caris Life Sciences

Fulfilling the promise of precision medicine through quality and innovation.

Senior DevOps Engineer, EKS/Kubernetes

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 1,001-5,000Since 2008H1B No SponsorCompany SiteLinkedIn

Location

United States

Posted

2 days ago

Salary

$122K - $148K / year

Seniority

Senior

Job Description

Senior DevOps Engineer, EKS/Kubernetes

Caris Life Sciences

• Design, deploy, and maintain Linux infrastructure in on-premises and cloud environments. • Automate infrastructure provisioning and configuration using tools such as Terraform, Ansible, or CloudFormation. • Manage and optimize AWS environments with a focus on performance, scalability, security, and cost efficiency. • Implement and maintain monitoring, logging, and alerting solutions (e.g., Datadog, Prometheus, Grafana, ELK, CloudWatch). • Architect, deploy, and operate production Kubernetes/AWS EKS clusters, including node group strategy, cluster upgrades, multi-tenant workload isolation, and cross-region disaster recovery (DR) architecture and build outs. • Define and lead cluster upgrade, security hardening, and disaster recovery strategies for production Kubernetes/AWS EKS environments at scale, while serving as a senior technical resource for complex production incidents. • Manage Kubernetes networking, including VPC CNI configuration and ingress controllers (ALB/NGINX/Traefik). • Implement IAM Roles for pod security standards, and network policies to secure EKS workloads. • Configure and tune cluster autoscaling (Cluster Autoscaler or Karpenter) and workload autoscaling (HPA/VPA) to optimize performance and cost. • Build and maintain Helm charts and GitOps-based deployment pipelines (e.g., ArgoCD, Flux) for Kubernetes workloads. • Manage Docker container builds and registries in support of EKS-based application deployment. • Deploy, scale, and maintain GitLab Runners (including Kubernetes executor runners on EKS) to support CI/CD pipeline throughput and reliability. • Support and help operate database platforms on AWS RDS (MySQL, PostgreSQL), collaborating with data owners on performance and reliability. • Ensure systems meet security and compliance requirements, including SOX and SOC 2 initiatives. • Execute and maintain Linux patching strategies, addressing security updates and CVEs in a timely manner. • Participate in incident response, root cause analysis, and recovery efforts. • Collaborate with development, QA, and cross-functional teams to improve reliability, release processes, and operational standards. • Participate in on-call rotations and provide after-hours support as required.

Job Requirements

  • Bachelor’s degree in computer science, Information Technology or related field
  • 8+ years of experience in Linux Systems Administration, DevOps, or Site Reliability Engineering roles
  • 5+ years of experience with AWS services, including EC2, VPC, IAM, RDS, S3, and CloudWatch
  • 5+ years of hands-on experience designing and operating production workloads on Kubernetes/AWS EKS, including cluster upgrades, networking, and autoscaling
  • Proficiency in scripting and automation using Python and Bash
  • Strong hands-on experience with Infrastructure as Code using Terraform, and with CI/CD pipelines (GitLab CI/CD), including running CI/CD workloads on Kubernetes/EKS
  • Proficiency with Docker, Helm, and Kubernetes troubleshooting in a production environment
  • Solid understanding of networking fundamentals and cloud security best practices
  • CKA (Certified Kubernetes Administrator) certification expected or actively in progress; Preferred Qualifications CKAD and AWS certifications (e.g., AWS Certified DevOps Engineer, Solutions Architect) a plus
  • Experience with Karpenter, Kyverno, OPA/Gatekeeper, Falco, and multi-cluster/multitenant EKS environments
  • Experience using AI and automation tools including Claude, Cursor, and OpenAI to streamline DevOps workflows through AI-assisted CI/CD, self-healing operations, and automated incident response
  • Experience with microservices, serverless architectures, and DevSecOps practices

Benefits

  • Highly competitive and inclusive medical, dental and vision coverage options
  • Health Savings Account for medical expenses and dependent care expenses
  • Flexible Spending Account to pay for certain out-of-pocket expenses
  • Paid time off, including: vacation, sick time and holidays
  • 401k match and Financial Planning tools
  • LTD and STD insurance coverages, as well as voluntary benefit options
  • Employee Assistance Program
  • Pet Insurance
  • Legal Assistance
  • Tuition Assistance

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Full TimeRemoteTeam 51-200H1B No Sponsor

• Own the cloud platform behind our engineering and AI ecosystem. • Design CI/CD pipelines, automate deployments, and improve security. • Manage Kubernetes environments and build Infrastructure as Code. • Enable engineers to deploy with confidence, collaborating with AI Architect and Engineering team.

India

AI-First SRE/DevOps Engineer

Axiad

Axiad is a cybersecurity company that provides enterprise-grade identity and access management (IAM) solutions, helping organizations securely manage credentials and authentication

DevOps Engineer2 days ago

Role Description Axiad is seeking a skilled AI-First SRE/DevOps Engineer with 5–8 years of hands-on infrastructure and platform engineering experience to help build and run Mesh, our Identity Visibility and Intelligence Platform (IVIP) — a cloud-native microservices platform on Kubernetes spanning human identity, non-human identity (NHI), post-quantum cryptography, and agentic AI identity risk. The ideal candidate has a builder mentality and a strong AI-First mindset: automation and AI are the default, not the afterthought, and infrastructure is something you create, not just maintain. This is a startup environment. You will own real surface area end-to-end, move fast, and ship. The role requires deep operational expertise in Kubernetes, CI/CD, and infrastructure-as-code, along with practical experience running AI/LLM systems in production. If your instinct when facing a repetitive task is to script it, agent-ify it, or delete it entirely — you'll fit right in. Responsibilities - Own reliability, observability, and delivery for a multi-tenant, cloud-native Kubernetes platform — from design through production, yours to run and yours to improve. - Build (not just operate) CI/CD pipelines, infrastructure-as-code, and GitOps-driven progressive delivery that let a small team ship many times a day, safely. - Embrace and advocate AI-First operations: automate incident response, runbooks, and remediation, and put AI agents in the loop to triage, diagnose, and propose fixes where it makes sense. Treat toil as a bug. - Build the infrastructure that AI-native features run on: inference gateways, LLM cost/latency observability, prompt/version pipelines, eval harnesses, and guardrails for agentic workloads. - Instrument everything — SLOs, error budgets, and distributed tracing across services and data pipelines. - Harden the platform: secrets management, supply-chain security, and least-privilege everywhere. - Troubleshoot and resolve production issues, leveraging AI-powered debugging and observability tooling. - Collaborate directly with product and platform engineers to translate requirements into resilient infrastructure — no throwing tickets over a wall; if you see a problem, it's yours to solve. - Mentor engineers in adopting AI-first operational practices and automation-by-default culture. Qualifications - 5–8 years of professional experience in SRE, DevOps, or platform engineering roles. - Builder mentality: you'd rather create a tool, platform, or automation than run a manual process twice. You ship things and stand behind them. - Ownership: you take problems from ambiguity to resolution without waiting for a ticket, a spec, or permission. When something you own breaks, you're the first to know and the first to act. - Strong Kubernetes operational experience — running it in production, not just deploying to it. - Demonstrable adoption of an AI-First mindset and tools (Claude Code, Cursor, or Windsurf). Daily use of at least one AI development tool is a must. - Fluency with infrastructure-as-code, GitOps, and modern CI/CD; comfortable scripting and building tooling (Go or Python preferred). - Cloud-native depth on at least one major cloud provider. - Solid observability expertise and SLO-driven operations experience. - Experience with containerization (Docker) and service mesh concepts. - Strong problem-solving skills and a collaborative mindset; excellent communication within Agile teams. - A bias for shipping — startup pace energizes you rather than stresses you. Preferred Qualifications - Experience building or operating LLM infrastructure: inference gateways, eval/observability tooling, agentic orchestration. - Data-pipeline and streaming/CDC experience. - Security or identity background; familiarity with post-quantum cryptography or supply-chain security. - Prior experience at an early-stage startup. Benefits - 120,000 - 160,000 OTE + Equity + Benefits

United States
$120K - $160K / year
ContractRemoteTeam 201-500Since 2014H1B No Sponsor

• Diseñar APIs/servicios backend que conecten y sincronicen datos entre AWS y Azure. • Construir conectores entre servicios cloud-nativos (ej. S3/SQS con Blob Storage/Event Grid). • Implementar arquitecturas orientadas a eventos (colas, webhooks, pub/sub) • Elegir el stack más adecuado por integración (Python, Node.js/TS, Java, Go u otros). • Diseñar contratos de API (REST/GraphQL/gRPC): versionado, auth (OAuth2/JWT), seguridad. • Escribir código probado y listo para producción; participar en code reviews. • Usar IA generativa (Copilot, Claude u otros) para acelerar desarrollo y documentación. • Diseñar/mantener pipelines CI/CD (GitHub Actions, ArgoCD) para los servicios que desarrolla. • Gestionar infraestructura como código (Terraform) en AWS y Azure. • Desplegar y operar en Kubernetes con Helm. • Dar observabilidad (Prometheus, Grafana, Datadog); rotaciones de guardia si aplica. • Explorar AIOps: IA/ML para detección de anomalías e incidentes. • Ser referente técnico ante el cliente en desarrollo y DevOps.

Mexico
Full TimeRemoteTeam 51-200Since 2003

• The RTE leads all Program Increment planning events, coordinating capacity, sequencing, and feature prioritization across the ART in alignment with program management, customer direction and COR guidance. • The RTE maintains the Integrated Master Schedule (IMS) on a weekly basis, manages release-level Project Process Agreements (PPAs) in accordance with FDA EPLC requirements, and ensures all sprint plans reflect realistic team capacity and accurate task-level estimates. • The RTE prepares and maintains release plan summaries, JIRA/ALM release plans mapped to sprint-level effort, and sprint status dashboards for review by the FDA Government IT PM. • The RTE serves as the primary integration point across all internal and external teams with dependencies on the program. • The RTE proactively identifies cross-team blockers, facilitates resolution, and escalates impediments to executive leadership when necessary. • The RTE drives participation from all impacted teams in end-to-end integration, regression, and UAT testing cycles for every FSDX release. • The RTE facilitates all ART-level ceremonies, including PI planning, System Demos, Inspect and Adapt workshops, and Scrum of Scrums. • The RTE coaches individual Scrum Masters and team-level Agile practices, fosters a culture of inspect and adapt, and drives measurable velocity improvements across the ART. • The RTE tracks and reports sprint burn rates, planned versus actual story point delivery, and release velocity trends, escalating potential overages to the Government IT PM immediately upon detection. • The RTE supports the Project Manager in preparing bi-weekly status reports, Monthly Status Reports (MSRs), and Monthly Financial Reports (MFRs) with accurate sprint metrics, burn rates, and risk indicators. • The RTE maintains the program-level Risk Management Plan, the RACI chart, and the Change Request Log, ensuring all items are tracked through closure and reflected in subsequent governance reporting. • The RTE maintains the program risk register and leads proactive risk identification, prioritization, and mitigation planning across all workstreams. • The RTE facilitates cross-team risk reviews during weekly status meetings and ensures newly identified risks are surfaced to FDA stakeholders immediately. • The RTE facilitates the integration of agency-approved AI tools and DevSecOps practices into the ART's delivery pipeline, supporting the program's commitment to leveraging AI for development efficiency, automated testing, intelligent workflow optimization, and documentation.

Maryland