Job Closed

This listing is no longer active.

Cisco logo
Cisco

We securely connect everything to make anything possible.

Customer Reliability Engineer – Hypershield

DevOps EngineerDevOps EngineerFull TimeRemoteLeadTeam 10,001+Since 1984H1B SponsorCompany SiteLinkedIn

Location

California

Posted

4 days ago

Salary

$158.2K - $200.7K / year

Seniority

Lead

Bachelor Degree8 yrs expEnglishLinux

Job Description

Customer Reliability Engineer – Hypershield

Cisco

• Own Hypershield cases escalated from Cisco TAC through to resolution • Diagnose complex production failures through the Hypershield surface • Localize faults across the layered architecture • Develop a deep understanding of each customer's architecture and configuration • Reproduce customer failures, partner with engineering to drive fixes • Convert individual cases into systemic improvements • Help build the team's proactive view of customer health

Job Requirements

  • Bachelor's + 8 years of experience, Master's + 6 years, or equivalent industry experience
  • Experience supporting enterprise customers in an escalation capacity
  • Experience operating and troubleshooting Cisco Nexus / NX-OS
  • Prior experience to localize failures across a layered data-center architecture
  • Linux operations experience at the command line

Benefits

  • medical, dental and vision insurance
  • a 401(k) plan with a Cisco matching contribution
  • paid parental leave
  • short and long-term disability coverage
  • basic life insurance
  • 10 paid holidays per full calendar year
  • 1 floating holiday for non-exempt employees
  • 1 paid day off for employee’s birthday
  • paid year-end holiday shutdown
  • 4 paid days off for personal wellness determined by Cisco
  • 16 days of paid vacation time per full calendar year
  • flexible vacation time off program
  • 80 hours of sick time off provided on hire date
  • additional paid time away may be requested to deal with critical issues
  • optional 10 paid days per full calendar year to volunteer

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Cisco logo

Lead Site Reliability Engineer, Engineering Enablement

Cisco

We securely connect everything to make anything possible.

DevOps Engineer4 days ago
Full TimeRemoteTeam 10,001+Since 1984H1B Sponsor

• architect, build, and evolve the developer experience for Meraki's Cloud Engineering teams • work within a team distributed across the US and UK and collaborate with other teams across SRE and Cloud Engineering • shape the day-to-day operations of Meraki's cloud • enable engineers to easily, safely, and confidently experiment on and build great products • lead the design and evolution of critical infrastructure for building, testing, and deploying our cloud applications • take the lead on complex problem resolution and debugging of internal and vendor-supplied tools • influence and drive operational excellence within the organization • learn and understand the priorities and practices of other engineering teams to design systems that work for them • partner with engineering leadership to define roadmaps, reporting on productivity gains and platform health to the SVP level • lead complex troubleshooting, perform blameless postmortems, and champion sustainable on-call practices across the organization

Massachusetts
$163.6K - $234.6K / year
Job Closed
OmegaHires logo

Cloud/DevOps Engineer

OmegaHires

Responsible recruiting!

DevOps Engineer4 days ago
ContractRemoteTeam 11-50H1B No Sponsor

• Designing, implementing, and maintaining cloud infrastructure • Ensuring the reliability and performance of applications • Working closely with development teams to establish CI/CD pipelines • Automating deployment processes • Contributing to critical projects from anywhere in the USA

United States
$90 - $100 / hour
Full TimeRemoteTeam 51-200

• Read, debug, and contribute to production C#/.NET code to diagnose and fix app-level reliability issues. • Identify and resolve memory leaks, thread pool exhaustion, and GC pressure before they manifest as incidents. • Partner with application engineers to embed reliability into new feature design and deployment practices. • Instrument .NET services with distributed tracing and structured logging to surface runtime anomalies early. • Operate and optimize EC2 Auto Scaling, ECS Fargate, and Lambda workloads — with clear judgment on when each is the right fit. • Build and maintain infrastructure-as-code using CloudFormation or CDK for consistent, reproducible environments. • Automate operational tasks, deployment pipelines, and disaster recovery procedures. • Continuously reduce toil through tooling and automation, freeing the team for higher-impact engineering work. • Manage RDS SQL Server deployments including Multi-AZ failover configuration and read replica setup. • Operate backup and point-in-time recovery (PITR) processes and validate restore procedures regularly. • Diagnose and resolve performance issues: slow queries, missing indexes, and blocking chains. • Capacity plan and scale database infrastructure to support transaction volume growth. • Build and maintain observability stacks using CloudWatch metrics, log insights, and alarms; AWS X-Ray for distributed tracing. • Own service health dashboards, SLOs/SLIs, and drive data-driven reliability improvements. • Design alerts that surface signal — not noise — and ensure on-call responders have the context to act quickly. • Conduct root cause analysis (RCA) on incidents and lead blameless post-mortems to capture lessons and prevent recurrence. • Design and maintain secure AWS network topologies: VPCs, subnets, security groups, and NACLs. • Configure and manage ALB/NLB routing, Route 53 DNS, and TLS certificate lifecycle via ACM. • Author and review least-privilege IAM policies; audit roles and resource-based policies for over-permissioning. • Support compliance and security controls relevant to a PCI-regulated payments environment. • Participate in on-call rotation to respond to production incidents and drive swift resolution. • Define and track error budgets; use them to balance velocity and reliability investment. • Communicate status updates clearly during incidents and coordinate cross-functional response. • Maintain and improve runbooks, escalation paths, and on-call health over time. • Collaborate with platform engineering teams on architecture decisions and scalability requirements. • Share observability and reliability best practices with application teams. • Mentor engineers on SRE principles and operational excellence.

United States
$145K - $160K / year
FluidStack logo

Principal Operations Engineer, Reliability

FluidStack

NVIDIA H100 & A100 GPUs available on demand at scale. Access thousands of GPUs for AI/LLM/ML, ready for deployment now.

DevOps Engineer4 days ago
Full TimeRemoteTeam 11-50H1B No Sponsor

• Own fleet reliability engineering: define availability targets, measure them honestly, and close the gap. • Run root cause analysis on the fleet's worst incidents and drive corrective actions to done across every site. • Build the failure data pipeline, facility and hardware both, that turns incident history into engineering priorities. • Set the maintenance strategy (reliability-centered, condition-based) so the fleet spends effort where the failure data says to.

United States
$220K - $260K / year