A leader in unified identity security
Staff Software Engineer – Reliability & Platform
Location
United Kingdom
Posted
124 days ago
Salary
0
Seniority
Lead
Job Description
Staff Software Engineer – Reliability & Platform
One Identity
• Own production operability by debugging complex issues, improving system visibility, and eliminating recurring problems at the source • Own production health for services — from detection through resolution to prevention • Improve mean time to detect (MTTD), mean time to resolve (MTTR), and recurrence rates for issues • Identify systemic issues and eliminate recurring problems through code fixes, architecture improvements, and better operational tooling • Improve observability across services — logs, metrics, and alerting — for faster diagnosis and resolution • Design and improve debugging workflows, runbooks, and internal tooling for engineers • Reduce operational burden by making systems easier to understand, operate, and troubleshoot • Partner closely with product teams to feed production learnings back into design and development • Reduce support and incident load by addressing root causes and improving system design, not just resolving individual issues
Job Requirements
- 4+ years of software engineering experience with ownership of production systems, reliability, or operational improvements
- Strong backend development experience (Ruby, Node.js, or similar)
- Solid understanding of REST APIs, service contracts, and software design principles
- Experience working across backend services, APIs, and production systems
- Experience building and operating services in AWS or similar cloud environments.
- Experience with observability, production debugging, and incident response.
- Willingness to participate in a mandatory 24/7 on-call rotation.
- Experience responding to production incidents and contributing to reliability improvements.
- Experience using, or strong interest in, AI-powered development tools (e.g., GitHub Copilot, ChatGPT, Cursor).
Benefits
- Competitive salary
- Flexible working hours
- Professional development budget
- Home office setup allowance
- Global team events
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Join Arista’s CloudVision-as-a-Service (CVaaS) global SRE team • Ensure scalability, reliability, and stability of global CloudVision service fleet • Develop, operate, and work with various databases • Contribute to automation and improvement of operational processes • Drive, develop, and lead projects in data platform architecture, capacity planning, and security
DevOps
AGtec Servicios InformáticosSomos especialistas en el desarrollo de soluciones informáticas para empresas de diversa índole.
• Desplegar infraestructura en la nube y OnPremises en Openshift y Docker. • Despliegue de Pipeline, revisión de servisores en ambientes productivos.
• Architect, implement, and continuously improve secure-by-design controls across multi-cloud environments • Develop and enforce Infrastructure as Code and policy-as-code guardrails • Design and maintain security controls within CI/CD pipelines • Lead threat modeling and architecture reviews • Define and promote secure coding standards
• Design and manage the architecture and operation of YPO's cloud infrastructure. • Lead the evaluation and adoption of new cloud services, platforms, and tooling. • Automate infrastructure management and provisioning. • Implement end-to-end CI/CD pipelines for mobile and backend applications. • Design and operate container orchestration infrastructure using Kubernetes. • Ensure observability and monitoring of systems to improve reliability and performance. • Collaborate with security engineers to integrate security into every stage of development and operations. • Mentor junior engineers and promote best practices within the engineering organization.



