Job Closed
This listing is no longer active.
Backblaze is the cloud storage innovator delivering a modern alternative to traditional cloud providers.
Site Reliability Engineer II
Location
India
Posted
123 days ago
Salary
0
Seniority
Mid Level
Job Description
Site Reliability Engineer II
Backblaze
• Support the availability and durability of critical services across production environments. • Monitor service health using SLIs, SLOs, and error budgets, and escalate issues when thresholds are at risk. • Participate in on-call rotations, incident response, and post-incident reviews to drive service improvements. • Follow established ITIL/OSS processes (incident, change, problem, and capacity management). • Develop automation for common operational tasks, reducing manual intervention and toil. • Contribute to monitoring, logging, and alerting frameworks (e.g., Prometheus, Grafana, Catchpoint,ELK). • Work with CI/CD pipelines, configuration management, and infrastructure as code tools (Terraform, Ansible, Jenkins). • Write scripts (Bash, Python, Go, etc.) to improve system reliability and efficiency. • Partner with engineering, product, and operations teams to support resilient system design and operations. • Assist in capacity planning and disaster recovery exercises. • Work with vendors and service providers to troubleshoot service issues and track SLA performance. • Document systems, share learnings, and help grow a reliability-minded engineering culture. • Contribute to playbooks, runbooks, and operational documentation. • Identify recurring issues and propose long-term improvements. • Promote reliability-focused practices within development and operations teams.
Job Requirements
- Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience).
- 2–4 years of experience in site reliability, systems engineering, or operations.
- Exposure to large-scale, production-grade systems.
- Solid Linux systems administration and troubleshooting skills.
- Familiarity with service reliability concepts - monitoring, alerting, incident response, and root cause analysis.
- Proficiency in at least one scripting language (Python, Bash, or Go).
- Understanding of containers (Kubernetes, Docker) and microservices concepts.
- Knowledge of incident response and operational best practices.
Benefits
- Paid time off
- Professional development opportunities
- Remote work options
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
DevOps Engineer
Time DoctorWorkforce analytics platform that gives managers actionable insights to improve team productivity and performance.
• Architecting, managing and scaling our cloud-native infrastructure on Google Cloud Platform and AWS • Work hands-on with modern serverless technologies, containerized architectures and container orchestration platforms • Leverage the full range of cloud services to ensure high availability, security and performance • Requires deep expertise in infrastructure such as code (Terraform), automated CI/CD pipelines and cloud-native best practices
Principal Site Reliability Engineer
Fidelity InvestmentsHere to help you grow your savings and invest in the life you want.
Title: Principal Site Reliability Engineer Location: 100 New Millennium Way, Bldg 1, Durham NC Job Description: Position Description: Combines Operational excellence with Development experience to deliver services at high scale, high availability with resilience. Builds reliability into the ecosystem by applying best practices in Resiliency Engineering, Automation, Observability and Chaos Testing. Streamlines and accelerates software delivery cycle by using DevOps practices and toolchain. Integrates Site Reliability Engineering (SRE) practices (Observability and Chaos) with DevOps processes and delivery pipelines to stop bad code from reaching production. Ensures business-critical enterprise systems are continuously available to internal and external customers. Implements technical standardization and process refinements within the engineering organization and for Site Reliability Engineers. Collaborates with production support teams to define and implement processes for the identification, collection, and analysis of incident data. Brings together technical, procedural, and financial data to reduce toil and increase efficiency. Primary Responsibilities: - Develops Chaos Testing capabilities using multiple Chaos Tools (AWS Fault Injection Service (FIS), Chaos Mesh, and Chaosd) and Chaos Toolkit. - Develops and enhances organization’s internal Chaos Framework to streamline Chaos Executions and reporting. - Provides specialized technical expertise in the adoption of Chaos Engineering by application teams. - Chaos tests and observes business-critical applications to understand the weaknesses and increase application resiliency. - Activates Observability for the critical applications with recommended Service Level Indicators and Service Level Objectives for Latency, Availability, Error Rate etc. - Utilizes modern monitoring tools (Datadog, Splunk, Catchpoint etc.) to reduce mean time to detect an issue and improve the response times. - Creates CI/CD pipelines with security and quality checks with Application Lifecycle management toolchain. Helps in integrating Chaos and Observability with CI/CD pipelines. - Automates repetitive activities using scripting languages (Python, Groovy etc.). - Implements and supports solutions based on cloud platforms AWS/Azure and container orchestration Kubernetes. - Onboards /Evaluates New Cloud services that help to enhance the Resiliency of cloud ecosystem. Serves as a liaison for vendor engagement. - Participates in incident management, problem management and incident postmortems. - Takes part in peer code reviews providing qualitative feedback. - Builds processes and capabilities to adapt and respond to risks, and disruptions, while maintaining business operations and data recovery with minimal disruptions. - Coaches peer SREs and application teams on SRE and DevOps. - Implements Agile methodologies in the team’s project completion using incremental and iterative steps. Education and Experience: Bachelor’s degree in Computer Science, Engineering, Information Technology, Information Systems, or a closely related field (or foreign education equivalent) and five (5) years of experience as a Principal Site Reliability Engineer (or closely related occupation) implementing resilient container and cloud-based applications and infrastructure solutions, using DevOps or SRE practices, in a financial services environment. Or, alternatively, Master’s degree (or foreign education equivalent) in Computer Science, Engineering, Information Technology, Information Systems, or a closely related field (or foreign education equivalent) and three (3) years of experience as a Principal Site Reliability Engineer (or closely related occupation) implementing resilient container and cloud-based applications and infrastructure solutions, using DevOps or SRE practices, in a financial services environment. Skills and Knowledge: Candidate must also possess: - Demonstrated Expertise (“DE”) improving application resiliency by implementing chaos engineering to build system's capability to withstand turbulent conditions in production, using Chaos Mesh, Chaosd, Azure Chaos Studio, AWS FIS, or Gremlin; and driving automation to implement scalable approaches for the planning, design, execution, and reporting of chaos testing using Jenkins pipelines, standard frameworks, data visualization, and dashboards. - DE implementing advanced observability practices and techniques in production and pre-production environments, at scale using Datadog, Splunk, or Catchpoint; tracking the error budget, proactively identifying issues, minimizing Mean Time to Repair (MTTR); and balancing customer expectations by implementing Service-Level Indicators (SLIs) and Service-Level Objectives (SLOs) using logs, traces, monitors and synthetic tests. - DE migrating and maintaining cloud applications and creating cloud solutions using Amazon Web Services (AWS) or Azure cloud services; Implementing infrastructure as code for cloud; Onboarding new AWS or Azure services with required reviews and security controls in non-production and production environments; and researching evolving cloud ecosystem to adopt machine learning based tools (AWS DevOps guru) to boost AIOps abilities. - DE implementing CI/CD pipelines in both production and non-production environments using Application Lifecycle Management (ALM) tools (JIRA, GitHub, Jenkins, SonarQube, Artifactory, or uDeploy) to enable faster code delivery, enhanced software quality, reliability, and security; and developing products, and core and common capabilities for the organization to reduce toil and drive standardization, using containerization and orchestration technologies (Docker or Kubernetes), Infrastructure as Code (IaC) tools, scripting languages (Python or Groovy), and engineering best practices. #PE1M2 #LI-DNI Certifications: Category:Information Technology Most roles at Fidelity are Hybrid, requiring associates to work onsite every other week (all business days, M-F) in a Fidelity office. This does not apply to Remote or fully Onsite roles. Some roles may have unique onsite requirements. Please consult with your recruiter for the specific expectations for this position. Please be advised that Fidelity’s business is governed by the provisions of the Securities Exchange Act of 1934, the Investment Advisers Act of 1940, the Investment Company Act of 1940, ERISA, numerous state laws governing securities, investment and retirement-related financial activities and the rules and regulations of numerous self-regulatory organizations, including FINRA, among others. Those laws and regulations may restrict Fidelity from hiring and/or associating with individuals with certain Criminal Histories.
Role Description The Senior Fullstack Engineer is a critical role responsible for developing core platform functionality. This includes creating tools to connect our Care Team with Members and building self-service options for Members. A key part of the role is defining the technical architecture and solutions necessary for efficient, rapid, and effective scaling. The goal is to deliver care to our Members that is easy, personalized, and highly effective. This role requires close collaboration with the engineering team, as well as with product and design, to successfully implement the defined product vision and roadmap. Key Responsibilities: - Delivery of full stack functionality on our solution that connects patients and clinicians through our web based portal and backend interfaces and APIs. - Support what is built, including monitoring, performance tuning, and responding to incidents. - Propose viable technical solutions to business needs that align with Ilant Health’s mission and values. - Contribute to the advisement of technical strategy, primarily related to architecting and scaling of current and new products. - Identify bottlenecks and implement improvements to processes, tools, and procedures. - Promote a culture of collaboration and learning across engineering, product, and design team via mentoring, documentation, presentations, or other knowledge sharing methods. Qualifications - Experience being on a small to medium sized engineering team (3 - 8 people) to deliver consumer or business facing features in a fast-paced environment. - Proven ability to deliver full stack development directly delivering value to patients and providers using technologies like Python, Next.JS and Typescript. - Effectively communicate between teams and within teams in order to drive alignment and increase effectiveness on delivery. - Ability to deal with ambiguity, demonstrate ownership, and lend your expertise to guiding the technical product roadmap. - Proven ability to switch domains and tech stacks then add value on Day 1. Requirements - Language: React, Python, Next.js, Typescript - Systems: AWS, Amplify, ECS, Postgres Benefits - Fully remote environment – work from anywhere while maintaining meaningful collaboration with a distributed team - Comprehensive health benefits – medical, dental, and vision coverage to support you and your family - Paid time off – 2 weeks of PTO to rest, recharge, and take the time you need - Flexible floating holiday – one additional day each year to celebrate what matters most to you - Paid sick leave – 5 sick days so you can prioritize your health when needed - 11 paid company holidays throughout the year - 401(k) retirement plan to help you invest in your future - Healthcare and Dependent Care FSA options for additional tax-advantaged savings
• Design and define system architecture for new or existing computer systems • Set-up, maintain, and develop continuous build/ integration infrastructure • Create and maintain fully automated CI build processes for multiple environments • Develop build and deployment scripts • Support CI/CD tools integration, operations, change management, and maintenance • Support full automation of CI/CD Development and Testing • Support policies, standards, guidelines, governance and related guidance for CI/CD operations and development • Enable successful release management by moving code from Development and Testing environments to Staging and Production



