Job Closed
This listing is no longer active.
Cisco is a publicly-traded, award-winning global technology solutions firm. Established in 1984 by a group of Stanford University computer scientists, Cisco has
Customer Reliability Engineer
Location
United States + 1 moreAll locations: United States | Canada
Posted
57 days ago
Salary
$158.2K - $241.8K / year
Seniority
Mid Level
No structured requirement data.
Job Description
Customer Reliability Engineer
Cisco
Role Description As a Customer Reliability Engineer (CRE), you are the tip of the spear in interacting with our customers. Our CRE team adapts the best practices of Site Reliability Engineering (SRE) and applies them to our customers. This role is focused on bringing this practice to the Hypershield software suite, running on Nexus Switches and as a standalone agent in various virtualization technologies, cloud providers and/or Kubernetes. - Gain a deep understanding of our customers and their architecture down into their various configurations. - Work with various stakeholders, internally and externally, to provide world-class support and issue resolution to various incidents. - Enhance our organization’s view into the health of our various customers. Qualifications - Bachelor’s + 8 years of experience or Master’s + 6 years of experience or equivalent industry experience, including experience in networking concepts and technologies across OSI layers 2 through 7. - 2+ years of direct experience supporting and engaging with enterprise customers in a technical capacity. - 1+ years of experience securing operating system (OS) instances, applications, and/or distributed systems. - 3+ years of hands-on experience operating Linux systems with experience in Kubernetes, cloud-native or container architecture. - 2+ years of hands-on experience configuring and managing Cisco Nexus switches in production environments. Requirements - Knowledge of standard methodologies for Linux operating systems security and their application in cloud-native technologies and environments. - Evidence of direct experience using network troubleshooting tools, including but not limited to packet capture and analysis utilities. - 2+ years of experience acting as a higher escalation point across multiple product lines. - Experience resolving issues with Kubernetes and cloud-native technologies in small to medium size Kubernetes environments. - Knowledge of Customer Reliability Engineering (CRE) practices, including Production Readiness Reviews (PRRs), Customer Test Environments (CuTEs), tooling, monitoring, knowledge base creation, and retrospectives. - Experience with at least one major cloud provider (AWS, Azure, or GCP). - Possessing CCNP Data Center, CCNP Enterprise, DevNet Professional, CCIE Data Center, CCIE Enterprise, or DevNet Expert certifications will be a key advantage. Benefits - Medical, dental, and vision insurance. - 401(k) plan with a Cisco matching contribution. - Paid parental leave. - Short and long-term disability coverage. - Basic life insurance. - 16 days of paid vacation time per full calendar year for non-exempt employees. - Flexible vacation time off program for exempt employees. - 80 hours of sick time off provided on hire date and each January 1st thereafter. - Optional 10 paid days per full calendar year to volunteer. - Employees may be eligible to receive grants of Cisco restricted stock units.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
DevOps Engineer – Red Hat OpenShift
EVOTEKToday’s Emerging Technology will be Tomorrow’s Competitive Advantage
• Design and develop solutions to complex application and integration challenges. • Oversight of our operations. • Leverage the latest technology in our cloud tenancies. • Cloud platform deployment hands-on experience in Azure and AWS. • Attend and actively participate in customer ceremonies and activities (Scrum/Kanban). • Creatively solve problems in the DevOps space, collaborating with customer Ops, Development, and QA team members. • Maintain a “can-do” attitude and a sense of urgency. • Listen to our customers/teams, understand their pain points, coach/mentor them for working smarter. • Document & Build CI/CD Pipelines. • Document & Build Infrastructure as Code. • Work with Docker and Kubernetes to create and schedule containers for deployments. • Document decisions regarding technology choices, best practices and process flow. • Automate builds and deployments across multi-platform environments.
Site Reliability Engineer II
BackblazeBackblaze is the cloud storage innovator delivering a modern alternative to traditional cloud providers.
• Support the availability and durability of critical services across production environments. • Monitor service health using SLIs, SLOs, and error budgets, and escalate issues when thresholds are at risk. • Participate in on-call rotations, incident response, and post-incident reviews to drive service improvements. • Follow established ITIL/OSS processes (incident, change, problem, and capacity management). • Develop automation for common operational tasks, reducing manual intervention and toil. • Contribute to monitoring, logging, and alerting frameworks (e.g., Prometheus, Grafana, Catchpoint,ELK). • Work with CI/CD pipelines, configuration management, and infrastructure as code tools (Terraform, Ansible, Jenkins). • Write scripts (Bash, Python, Go, etc.) to improve system reliability and efficiency. • Partner with engineering, product, and operations teams to support resilient system design and operations. • Assist in capacity planning and disaster recovery exercises. • Work with vendors and service providers to troubleshoot service issues and track SLA performance. • Document systems, share learnings, and help grow a reliability-minded engineering culture. • Contribute to playbooks, runbooks, and operational documentation. • Identify recurring issues and propose long-term improvements. • Promote reliability-focused practices within development and operations teams.
• Fleet Management at Scale: Design, implement, and maintain robust and secure device management strategies for remote devices using Unified Endpoint Management (UEM), MDM solutions, and orchestration tools. • Reliability & Monitoring: Develop and manage observability pipelines to track device health, connectivity, and performance metrics across diverse warehouse environments. • OTA & Lifecycle Management: Own the end-to-end lifecycle of device software, including secure Over-the-Air (OTA) firmware updates, rollback strategies, and OS hardening. • Incident Response: Participate in on-call rotations to troubleshoot complex system failures, performing root cause analysis (RCA) to drive long-term reliability improvements. • Self-Healing Infrastructure: Develop automated remediation scripts that detect and fix common edge issues such as hung scanning processes or display driver freezes without manual intervention. • Zero-Touch Scalability: Architect and maintain remote provisioning and management workflows for a global fleet of Linux, iPads, and Android devices using secure remote management strategies. • Secure Remote Access: Implement and manage secure remote access protocols such as SSH, VPNs, and private APNs to enable out-of-band troubleshooting and real-time device control without physical site visits. • SLO/SLI Frameworks: Define and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for device availability, connectivity, and peripheral performance. • Error Budget Management: Use error budgets to balance the pace of innovation with fleet reliability, ensuring data-driven decisions for feature releases versus stability fixes. • Security Governance: Align fleet operations with industry standards such as the NIST Cybersecurity Framework (CSF), ISO/IEC 27001, and CIS Controls. • Vulnerability Management: Drive continuous monitoring and automated patching schedules to mitigate risks and ensure regulatory compliance across all managed device platforms.
Role Description As the Staff DevSecOps Engineer, you will be the technical owner of how security is built into Trase's software development lifecycle and cloud operations. - Integrate automated security testing, continuous vulnerability management, and secure coding practices directly into existing CI/CD pipelines. - Own the implementation of Trase's dedicated security architecture, delivering shift-left tooling (SAST, DAST, SCA, secrets scanning, and IaC scanning) alongside production cloud security services. - Standardize and operate secure pipelines to empower Trase's software engineers while maintaining required controls and capabilities. Qualifications - 10+ years of experience in security engineering, DevSecOps, cloud security, or platform security roles. - Deep, hands-on experience securing modern CI/CD pipelines. - Strong cloud security expertise, primarily in Google Cloud Platform. - Expert-level Terraform skills with a track record of building secure-by-default IaC modules. - Demonstrated experience with SIEM operations and incident response leadership. - Practical experience in environments governed by SOC 2, HIPAA, and ISO 27001. - Strong programming or scripting skills (Python, Go, or similar). - Excellent partnership skills and a developer-empathetic mindset. - Strong affinity for working with LLMs and AI agents. - US Citizen and eligible for US security clearance. Requirements - Design, implement, and operate the shift-left security toolchain across Trase's CI/CD pipelines. - Define how findings are triaged, routed, and remediated. - Establish and enforce policy-as-code and pre-merge security gates. - Design and deploy Trase's production cloud security architecture. - Implement foundational controls including network segmentation and workload identity. - Build, codify, and maintain secure-by-default infrastructure modules in Terraform. - Operate and fine-tune Trase's SIEM and security telemetry pipeline. - Enhance and lead aspects of Trase's technical security incident response capability. - Operate the end-to-end vulnerability management lifecycle. - Partner closely with Engineering and the broader Security and Compliance team. - Mentor junior Security and Compliance engineers and members of the Engineering team. Benefits - Career track opportunity with potential for rapid advancement. - 100% employer paid, comprehensive health care including medical, dental, and vision for you and your family. - Paid maternity and paternity for 14 weeks at employees' normal pay. - Unlimited PTO, with management approval. - Opportunities for professional development and continued learning. - Optional 401K, FSA, and equity incentives available. - Mental health benefits available through Tara Mind. - Cost effective GLP-1 solutions available through Crux.




