Job Closed
This listing is no longer active.
Set data in motion.
Staff Software Engineer I - SRE
Location
India
Posted
56 days ago
Salary
0
Seniority
Lead
No structured requirement data.
Job Description
Staff Software Engineer I - SRE
Confluent
Role Description Confluent Cloud processes millions of events per second across AWS, GCP, and Azure. When incidents happen in a multi-cloud streaming platform, they happen at scale—data in motion, exactly-once semantics, and cascading failure modes that require deep systems thinking. We need an expert-level engineer who can drive proactive reliability improvements that prevent these incidents before they occur. This role combines hands-on technical work with strategic program ownership. You'll spend roughly 75% of your time on engineering: - Building automation - Improving tooling - Analyzing systemic failure patterns - Designing reliability improvements The remaining 25% is teaching and coordination: - Coaching teams through post-mortems - Training incident commanders - Evolving our incident response practices You'll be part of a global team with follow-the-sun coverage, with clean handoffs that keep everyone working sustainable hours. Confluent has 800-1000 engineers across highly autonomous teams. This role sits within Cloud Architecture and Reliability - Supportability (CAR-S), a horizontal team that owns reliability standards and tooling across engineering. You're the person who makes us need incident management less. What You Will Do - Proactive Reliability Engineering (~75% of role) - Analyze systemic failure patterns and design improvements that prevent incident recurrence - Define and maintain SLO/SLA frameworks; use error budgets to guide reliability investments - Build tooling and automation to reduce incident response toil and scale team impact - Own Rootly configuration, workflows, and integrations with PagerDuty, Jira, Confluence, and Slack - Analyze reliability data to identify systemic improvements; build dashboards that drive action - Explore AI-assisted approaches to documentation quality and incident analysis - Design scalable reliability standards that reduce reactive workload over time - Incident Management Program (~25% of role) - Own standards, practices, and continuous improvement of incident response - Serve as an on-call Incident Commander for production incidents, including acting as escalation IC when incidents exceed a team's management chain - Develop and deliver training programs for engineering teams at all levels - Coach teams through post-mortems and on developing actionable corrective actions - Customer Root Cause Analysis (CRCA) - Edit and review customer-facing incident documents to ensure quality and clarity - Drive turnaround SLAs while maintaining technical accuracy - Ensure clear explanation of what happened, why, and how we'll prevent recurrence - Cross-Team Leadership - Partner with engineering leaders to elevate reliability practices - Be the expert who teams proactively engage for guidance Qualifications - 10+ years in SRE, incident management, or reliability engineering - Cloud experience with at least one of AWS, GCP, or Azure - Deep expertise with incident management tooling (Rootly, PagerDuty, or similar platforms) - Strong understanding of distributed systems and failure modes at scale—Kafka/event streaming expertise preferred, or demonstrated rapid mastery of complex systems - Deep experience with observability: metrics, logging, tracing—ability to diagnose complex issues - Kubernetes and container orchestration experience - Understanding of CI/CD pipelines and release processes - Systems thinking: understanding how infrastructure design choices affect failure modes and recovery - Familiarity with SLO/SLA frameworks - Track record as a trusted advisor across engineering organizations - Experience driving org-wide process and cultural changes - Strong written communication (design docs, one-pagers, runbooks) - Post-mortem facilitation experience - Experience with async collaboration across time zones - Large company experience navigating reliability/incident programs at 500+ engineer organizations What Gives You an Edge - Multi-cloud experience (minimum 2+ of AWS/GCP/Azure) - Modern CI/CD, GitHub, AI-assisted workflows—you'll have the freedom to build what you need Benefits - Belonging isn’t a perk here. It’s the baseline. - We work across time zones and backgrounds, knowing the best ideas come from different perspectives. - We make space for everyone to lead, grow, and challenge what’s possible. - We’re proud to be an equal opportunity workplace.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Collaborate closely with fellow devops engineers and the development team to deploy and maintain application infrastructure. • Assist in the development and support of tooling to streamline the deployment and maintenance of our products. • Work with Github, Jenkins, and Chef to deploy applications from development through to production environments. • Support both in-house and third-party applications, including handling deployments, upgrades, and troubleshooting. • Build and manage automation pipelines for application deployment and maintenance. • Engage in the day-to-day management of Linux servers via the command line. • Create monitoring dashboards and alerts in Grafana leveraging Prometheus and Alertmanager. • Document processes and best practices clearly and concisely. • Participate in incident solving on-call rotation
• Own end-to-end release and deployment lifecycle: build → package → deploy → verify → rollback • Develop and support **Octopus Deploy** projects, lifecycles, channels, variables, and deployment processes • Implement deployment automation with **Ansible** (playbooks/roles, inventories, idempotent changes) • Maintain Git-based release workflows in **GitHub** (branching, tagging, versioning, release notes) • Build/maintain CI pipelines in GitHub Actions (or existing tooling) to produce artifacts and trigger Octopus releases • Standardize deployment patterns across applications (templates, shared steps, reusable Ansible roles) • Manage environment configuration and secrets in a controlled way (variable sets, permissions, auditing) • Improve deployment safety: approvals, health checks, smoke tests, automated validation, and rollback strategies • Support production releases, troubleshoot deployment failures, and drive root-cause analysis • Maintain release documentation, runbooks, and change management practices • Collaborate with developers, QA, and operations to plan releases and reduce downtime
• Act as technical lead for DevOps/Platform/Release engineering: set direction, standards, and best practices • Architect and govern end-to-end delivery: infrastructure provisioning, configuration management, CI/CD, release processes, and operations • Design and support Windows-based high availability solutions, with deep ownership of Windows clustering (failover/HA patterns, maintenance, upgrades, troubleshooting) • Lead Linux automation and platform standardization (configuration, patching, hardening, performance tuning) • Own Infrastructure as Code strategy with Terraform (modules, environments, state, governance) • Own automation strategy with Ansible (reusable roles, inventories, secure secrets handling, idempotency) • Build and standardize deployments using Octopus Deploy, GitHub, and Ansible (templates, shared steps, release promotion, rollback) • Design and mature CI/CD pipelines (artifact versioning, approvals, promotion strategy, policy-as-code where applicable) • Establish observability standards using VictoriaMetrics/Prometheus (metrics strategy, alerting, SLO/SLA monitoring, dashboards) • Provide production leadership: incident response, RCA/postmortems, reliability improvements, capacity planning • Mentor engineers, review designs/code, and raise overall engineering quality across teams • Produce and maintain architecture docs, runbooks, and platform roadmaps
• Build and maintain CI/CD pipelines for application builds, automated testing, packaging, and deployment activities. • Implement automation solutions for environment provisioning, operational workflows, release processes, and infrastructure support tasks. • Support secure delivery practices including code scanning, dependency validation, secrets management, and policy enforcement activities. • Troubleshoot and resolve build, deployment, pipeline, and environment-related issues across multiple applications and services. • Collaborate with development and QA teams to improve release quality, deployment reliability, and software delivery timelines. • Support cloud-based infrastructure and shared platform services in coordination with engineers, architects, and operations teams. • Maintain documentation for deployment pipelines, environment configurations, release procedures, and operational support processes. • Participate in incident response efforts, root cause analysis, and continuous process improvement initiatives. • Monitor system and pipeline performance and recommend improvements to automation, tooling, and workflow efficiency. • Support change management, deployment coordination, and release readiness activities across production and non-production environments. • Contribute to various projects and initiatives as assigned, demonstrating adaptability and a collaborative mindset.


