Job Closed
This listing is no longer active.
We securely connect everything to make anything possible.
Lead Site Reliability Engineer, Engineering Enablement
Location
Massachusetts
Posted
5 days ago
Salary
$163.6K - $234.6K / year
Seniority
Senior
Job Description
Lead Site Reliability Engineer, Engineering Enablement
Cisco
• architect, build, and evolve the developer experience for Meraki's Cloud Engineering teams • work within a team distributed across the US and UK and collaborate with other teams across SRE and Cloud Engineering • shape the day-to-day operations of Meraki's cloud • enable engineers to easily, safely, and confidently experiment on and build great products • lead the design and evolution of critical infrastructure for building, testing, and deploying our cloud applications • take the lead on complex problem resolution and debugging of internal and vendor-supplied tools • influence and drive operational excellence within the organization • learn and understand the priorities and practices of other engineering teams to design systems that work for them • partner with engineering leadership to define roadmaps, reporting on productivity gains and platform health to the SVP level • lead complex troubleshooting, perform blameless postmortems, and champion sustainable on-call practices across the organization
Job Requirements
- Bachelor's degree and 8+ years of relevant experience, or Master's degree and 6+ years of relevant experience or equivalent related work experience
- Have a software development background with 5+ years experience coding with languages like Ruby or Python.
- Rapid technical adoption experience, evaluating and adopting new technologies with quick turnaround to address critical production gaps or security vulnerabilities.
- Experience mentoring and providing technical oversight for a team
- Experience with systems at scale, managing and operating distributed systems at a scale of 1,000+ nodes or high-concurrency environments with an owner/operator mindset.
- Automation experience in configuration-as-code and infrastructure automation.
- Have experience with modern Unix/Linux operating systems/distributions.
Benefits
- medical, dental and vision insurance
- a 401(k) plan with a Cisco matching contribution
- paid parental leave
- short and long-term disability coverage
- basic life insurance
- 10 paid holidays per full calendar year
- 1 floating holiday for non-exempt employees
- 1 paid day off for employee’s birthday
- paid year-end holiday shutdown
- 4 paid days off for personal wellness determined by Cisco
- Non-exempt employees receive 16 days of paid vacation time per full calendar year
- Exempt employees participate in Cisco’s flexible vacation time off program
- 80 hours of sick time off provided on hire date and each January 1st thereafter
- up to 80 hours of unused sick time carried forward from one calendar year to the next
- Optional 10 paid days per full calendar year to volunteer
- Employees are also eligible to earn annual bonuses subject to Cisco’s policies.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Designing, implementing, and maintaining cloud infrastructure • Ensuring the reliability and performance of applications • Working closely with development teams to establish CI/CD pipelines • Automating deployment processes • Contributing to critical projects from anywhere in the USA
• Read, debug, and contribute to production C#/.NET code to diagnose and fix app-level reliability issues. • Identify and resolve memory leaks, thread pool exhaustion, and GC pressure before they manifest as incidents. • Partner with application engineers to embed reliability into new feature design and deployment practices. • Instrument .NET services with distributed tracing and structured logging to surface runtime anomalies early. • Operate and optimize EC2 Auto Scaling, ECS Fargate, and Lambda workloads — with clear judgment on when each is the right fit. • Build and maintain infrastructure-as-code using CloudFormation or CDK for consistent, reproducible environments. • Automate operational tasks, deployment pipelines, and disaster recovery procedures. • Continuously reduce toil through tooling and automation, freeing the team for higher-impact engineering work. • Manage RDS SQL Server deployments including Multi-AZ failover configuration and read replica setup. • Operate backup and point-in-time recovery (PITR) processes and validate restore procedures regularly. • Diagnose and resolve performance issues: slow queries, missing indexes, and blocking chains. • Capacity plan and scale database infrastructure to support transaction volume growth. • Build and maintain observability stacks using CloudWatch metrics, log insights, and alarms; AWS X-Ray for distributed tracing. • Own service health dashboards, SLOs/SLIs, and drive data-driven reliability improvements. • Design alerts that surface signal — not noise — and ensure on-call responders have the context to act quickly. • Conduct root cause analysis (RCA) on incidents and lead blameless post-mortems to capture lessons and prevent recurrence. • Design and maintain secure AWS network topologies: VPCs, subnets, security groups, and NACLs. • Configure and manage ALB/NLB routing, Route 53 DNS, and TLS certificate lifecycle via ACM. • Author and review least-privilege IAM policies; audit roles and resource-based policies for over-permissioning. • Support compliance and security controls relevant to a PCI-regulated payments environment. • Participate in on-call rotation to respond to production incidents and drive swift resolution. • Define and track error budgets; use them to balance velocity and reliability investment. • Communicate status updates clearly during incidents and coordinate cross-functional response. • Maintain and improve runbooks, escalation paths, and on-call health over time. • Collaborate with platform engineering teams on architecture decisions and scalability requirements. • Share observability and reliability best practices with application teams. • Mentor engineers on SRE principles and operational excellence.
Principal Operations Engineer, Reliability
FluidStackNVIDIA H100 & A100 GPUs available on demand at scale. Access thousands of GPUs for AI/LLM/ML, ready for deployment now.
• Own fleet reliability engineering: define availability targets, measure them honestly, and close the gap. • Run root cause analysis on the fleet's worst incidents and drive corrective actions to done across every site. • Build the failure data pipeline, facility and hardware both, that turns incident history into engineering priorities. • Set the maintenance strategy (reliability-centered, condition-based) so the fleet spends effort where the failure data says to.
Site Reliability Engineer, Big Data
PulsePointWebMD and its affiliates is an Equal Opportunity/Affirmative Action employer and does not discriminate on the basis of race, ancestry, color, religion, sex, gender, age, marital status, sexual orientation, gender identity, national origin, medical condition, disability, veterans status, or any other basis protected by law.
Role Description Build streaming and storage systems at PulsePoint. PulsePoint processes billions of events daily through Kafka, Hadoop, HDFS, and Ceph. Our Data Platform team maintains these systems across hybrid infrastructure: bare-metal on-prem, cloud, and the integration between them. You own the lifecycle from architecture through deployment to capacity planning and incident response. What you'll work on: - Kafka architecture, topic design, governance, partition strategy, throughput and latency optimization. - Ceph operations, pool design, placement optimization, capacity planning. - Operational automation, reduce manual work, faster incident response, preventive systems. - SQL Server backup and recovery pipelines, basic cluster support. - Data team tooling with self-service capabilities and observability. Qualifications - You've operated data infrastructure at petabyte scale or billions of events/day. - You understand replication failures, consistency tradeoffs, and cost-optimization of large systems. - You automate before troubleshooting. - You take ownership across system layers, not just one component. - You simplify complex systems. Requirements - 5+ years operating distributed systems at scale in production. - Deep expertise in Kafka, Ceph, or similar distributed infrastructure. - Proven ability to design for scale and reliability. - Experience mentoring engineers and making technical decisions. - Willingness to work 9am-6pm ET U.S. hours. Benefits - This is not a ticket-driven operational role. - You'll help define platform architecture, influence engineering standards and work on infrastructure that supports multiple engineering organizations. - The engineer joining this role is expected to become a key technical contributor shaping the future of the platform. - You can work fully remotely and get the opportunity to work on large-scale infrastructure and distributed systems challenges. Hiring Process - Introductory conversation (~60m): Learn about your background and discuss the role. - Technical discussion (~60m): Deep dive into systems engineering, Kubernetes and operational experience. - Architecture discussion (~60m): Explore platform design, distributed systems and technical decision-making. - Leadership conversation (~30m): Meet engineering leadership and discuss team, strategy and long-term direction.




