Symbotic logo
Symbotic

Reinvent the warehouse®. Reimagine the supply chain®.

Senior Reliability Engineer

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 501-1,000Since 2007H1B SponsorCompany SiteLinkedIn

Location

United States

Posted

4 days ago

Salary

$120K - $165K / year

Seniority

Senior

Bachelor Degree8 yrs expEnglishDistributed Systems

Job Description

Senior Reliability Engineer

Symbotic

• Lead high-impact RCA investigations for complex production incidents spanning software, infrastructure, industrial controls, and production SOP execution • Chair structured, blameless RCA reviews aligned with ITIL Problem Management — fact-based analysis, clear ownership, timely resolution • Serve as the customer-facing technical lead for RCA discussions, updates, and formal deliverables within defined SLA timelines • Analyze logs, telemetry, and incident trends across many issues to find recurring failure patterns — then influence teams across the organization to eliminate them • Present findings, risks, and recommendations to senior internal leadership and customer stakeholders, backed by data you own end to end • Drive Continuous Service Improvement initiatives that measurably reduce repeat incidents and investigation toil through automation, tooling, and reporting • Mentor teammates and elevate RCA quality standards as a senior individual contributor

Job Requirements

  • Minimum 8 years supporting complex, business-critical production environments, with a career centered on reliability
  • Minimum 5 years leading technical RCA, post-incident reviews, or ITIL-aligned Problem Management across software, infrastructure, systems, or industrial technology domains
  • Strong hands-on troubleshooting and data analysis across large-scale distributed systems, on-prem infrastructure, custom software, logs, telemetry, and incident datasets
  • A track record of regular, ongoing customer interaction — you can tell us who you worked with, at what level, and how often — and of earning trust with executives, technical and non-technical alike
  • Proven ability to run multiple high-priority investigations in parallel while influencing cross-functional teams, without direct authority, to close actions on time
  • Bachelor’s degree in a technical field, or equivalent practical experience

Benefits

  • medical
  • dental
  • vision
  • disability
  • 401K
  • PTO

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Caris Life Sciences logo

Senior DevOps Engineer, EKS/Kubernetes

Caris Life Sciences

Fulfilling the promise of precision medicine through quality and innovation.

DevOps Engineer4 days ago
Full TimeRemoteTeam 1,001-5,000Since 2008H1B No Sponsor

• Design, deploy, and maintain Linux infrastructure in on-premises and cloud environments. • Automate infrastructure provisioning and configuration using tools such as Terraform, Ansible, or CloudFormation. • Manage and optimize AWS environments with a focus on performance, scalability, security, and cost efficiency. • Implement and maintain monitoring, logging, and alerting solutions (e.g., Datadog, Prometheus, Grafana, ELK, CloudWatch). • Architect, deploy, and operate production Kubernetes/AWS EKS clusters, including node group strategy, cluster upgrades, multi-tenant workload isolation, and cross-region disaster recovery (DR) architecture and build outs. • Define and lead cluster upgrade, security hardening, and disaster recovery strategies for production Kubernetes/AWS EKS environments at scale, while serving as a senior technical resource for complex production incidents. • Manage Kubernetes networking, including VPC CNI configuration and ingress controllers (ALB/NGINX/Traefik). • Implement IAM Roles for pod security standards, and network policies to secure EKS workloads. • Configure and tune cluster autoscaling (Cluster Autoscaler or Karpenter) and workload autoscaling (HPA/VPA) to optimize performance and cost. • Build and maintain Helm charts and GitOps-based deployment pipelines (e.g., ArgoCD, Flux) for Kubernetes workloads. • Manage Docker container builds and registries in support of EKS-based application deployment. • Deploy, scale, and maintain GitLab Runners (including Kubernetes executor runners on EKS) to support CI/CD pipeline throughput and reliability. • Support and help operate database platforms on AWS RDS (MySQL, PostgreSQL), collaborating with data owners on performance and reliability. • Ensure systems meet security and compliance requirements, including SOX and SOC 2 initiatives. • Execute and maintain Linux patching strategies, addressing security updates and CVEs in a timely manner. • Participate in incident response, root cause analysis, and recovery efforts. • Collaborate with development, QA, and cross-functional teams to improve reliability, release processes, and operational standards. • Participate in on-call rotations and provide after-hours support as required.

United States
$122K - $148K / year
Full TimeRemoteTeam 51-200H1B No Sponsor

• Own the cloud platform behind our engineering and AI ecosystem. • Design CI/CD pipelines, automate deployments, and improve security. • Manage Kubernetes environments and build Infrastructure as Code. • Enable engineers to deploy with confidence, collaborating with AI Architect and Engineering team.

India

AI-First SRE/DevOps Engineer

Axiad

Axiad is a cybersecurity company that provides enterprise-grade identity and access management (IAM) solutions, helping organizations securely manage credentials and authentication

DevOps Engineer4 days ago

Role Description Axiad is seeking a skilled AI-First SRE/DevOps Engineer with 5–8 years of hands-on infrastructure and platform engineering experience to help build and run Mesh, our Identity Visibility and Intelligence Platform (IVIP) — a cloud-native microservices platform on Kubernetes spanning human identity, non-human identity (NHI), post-quantum cryptography, and agentic AI identity risk. The ideal candidate has a builder mentality and a strong AI-First mindset: automation and AI are the default, not the afterthought, and infrastructure is something you create, not just maintain. This is a startup environment. You will own real surface area end-to-end, move fast, and ship. The role requires deep operational expertise in Kubernetes, CI/CD, and infrastructure-as-code, along with practical experience running AI/LLM systems in production. If your instinct when facing a repetitive task is to script it, agent-ify it, or delete it entirely — you'll fit right in. Responsibilities - Own reliability, observability, and delivery for a multi-tenant, cloud-native Kubernetes platform — from design through production, yours to run and yours to improve. - Build (not just operate) CI/CD pipelines, infrastructure-as-code, and GitOps-driven progressive delivery that let a small team ship many times a day, safely. - Embrace and advocate AI-First operations: automate incident response, runbooks, and remediation, and put AI agents in the loop to triage, diagnose, and propose fixes where it makes sense. Treat toil as a bug. - Build the infrastructure that AI-native features run on: inference gateways, LLM cost/latency observability, prompt/version pipelines, eval harnesses, and guardrails for agentic workloads. - Instrument everything — SLOs, error budgets, and distributed tracing across services and data pipelines. - Harden the platform: secrets management, supply-chain security, and least-privilege everywhere. - Troubleshoot and resolve production issues, leveraging AI-powered debugging and observability tooling. - Collaborate directly with product and platform engineers to translate requirements into resilient infrastructure — no throwing tickets over a wall; if you see a problem, it's yours to solve. - Mentor engineers in adopting AI-first operational practices and automation-by-default culture. Qualifications - 5–8 years of professional experience in SRE, DevOps, or platform engineering roles. - Builder mentality: you'd rather create a tool, platform, or automation than run a manual process twice. You ship things and stand behind them. - Ownership: you take problems from ambiguity to resolution without waiting for a ticket, a spec, or permission. When something you own breaks, you're the first to know and the first to act. - Strong Kubernetes operational experience — running it in production, not just deploying to it. - Demonstrable adoption of an AI-First mindset and tools (Claude Code, Cursor, or Windsurf). Daily use of at least one AI development tool is a must. - Fluency with infrastructure-as-code, GitOps, and modern CI/CD; comfortable scripting and building tooling (Go or Python preferred). - Cloud-native depth on at least one major cloud provider. - Solid observability expertise and SLO-driven operations experience. - Experience with containerization (Docker) and service mesh concepts. - Strong problem-solving skills and a collaborative mindset; excellent communication within Agile teams. - A bias for shipping — startup pace energizes you rather than stresses you. Preferred Qualifications - Experience building or operating LLM infrastructure: inference gateways, eval/observability tooling, agentic orchestration. - Data-pipeline and streaming/CDC experience. - Security or identity background; familiarity with post-quantum cryptography or supply-chain security. - Prior experience at an early-stage startup. Benefits - 120,000 - 160,000 OTE + Equity + Benefits

United States
$120K - $160K / year
ContractRemoteTeam 201-500Since 2014H1B No Sponsor

• Diseñar APIs/servicios backend que conecten y sincronicen datos entre AWS y Azure. • Construir conectores entre servicios cloud-nativos (ej. S3/SQS con Blob Storage/Event Grid). • Implementar arquitecturas orientadas a eventos (colas, webhooks, pub/sub) • Elegir el stack más adecuado por integración (Python, Node.js/TS, Java, Go u otros). • Diseñar contratos de API (REST/GraphQL/gRPC): versionado, auth (OAuth2/JWT), seguridad. • Escribir código probado y listo para producción; participar en code reviews. • Usar IA generativa (Copilot, Claude u otros) para acelerar desarrollo y documentación. • Diseñar/mantener pipelines CI/CD (GitHub Actions, ArgoCD) para los servicios que desarrolla. • Gestionar infraestructura como código (Terraform) en AWS y Azure. • Desplegar y operar en Kubernetes con Helm. • Dar observabilidad (Prometheus, Grafana, Datadog); rotaciones de guardia si aplica. • Explorar AIOps: IA/ML para detección de anomalías e incidentes. • Ser referente técnico ante el cliente en desarrollo y DevOps.

Mexico