We built Stories so you don't have to.
Site Reliability Engineer
Location
Tunisia
Posted
2 days ago
Salary
€25K / year
Seniority
Senior
Job Description
Site Reliability Engineer
Storyteller
• Respond to live incidents • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state. • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code. • Use AI throughout triage and diagnosis while checking its conclusions against real evidence. • Choose and execute a proportionate mitigation, rollback, repair or bounded fix. • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green. • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication. • Join customer conversations occasionally when direct technical involvement is genuinely useful. • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix. • Escalate with evidence, customer impact, actions already taken and the specific decision or help required. • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait. • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners. • Remove, consolidate and tune low-value alerts, and design monitoring around real service and customer outcomes. • Analyse material incidents with AI, validate the conclusions and turn repeated failure patterns into better alerts, runbooks, AI Skills, automation or product improvements. • Improve dashboards, diagnostics, service ownership and escalation information so common incidents are easier to understand and resolve. • Create safe, supervised automation for common operational actions. • Work with product teams to close observability, rollback, runbook and supportability gaps. • Detect and help contain unusual service-cost behaviour, then route wider follow-up to the appropriate cost or product owner. • Make reliability and on-call performance easier for the company to understand and improve over time.
Job Requirements
- Previous responsibility for live production systems or an on-call rota is strongly preferred because it is useful evidence that you understand the realities of incident response.
- Operational judgement - You can separate customer impact, symptoms and likely causes, make practical decisions under uncertainty and recognise when an intervention is no longer safe or bounded.
- Technical comfort and aptitude - You are comfortable exploring unfamiliar systems through code, logs, APIs, data, infrastructure and command-line tools, and can make hands-on changes with a clear validation plan.
- AI-native execution - You use AI for substantive technical work - investigation, hypothesis generation, code, automation, incident analysis and workflow improvement - while supervising the agent and challenging its conclusions.
- Accuracy and validation discipline - You actively look for false confidence and verify outcomes through appropriate technical and customer signals.
- Systems thinking - You look for repeated patterns and improve the triggers, owners, runbooks, automation, metrics and feedback loops around the work.
- Clear coordination and communication - You communicate calmly and concisely with Support, developers and non-technical stakeholders, making evidence, impact, uncertainty, ownership and next actions easy to understand.
- Curiosity and resilience - You learn unfamiliar products and tools quickly, keep investigating when the first hypothesis fails and change your approach when the evidence demands it.
- Nice to have Cloud platforms such as Azure or Cloudflare.
- Distributed application and API diagnostics.
- Databases, queues and background-processing systems.
- Observability, alerting and incident-management platforms.
- Infrastructure, deployment and release automation.
- Application development and safe production debugging.
- AI coding agents and workflow automation.
Benefits
- Fully remote working from anywhere in Tunisia!
Related Guides
Related Categories
Related Job Pages
More Production Engineer Jobs
• Respond to live incidents • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state. • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code. • Use AI throughout triage and diagnosis while checking its conclusions against real evidence. • Choose and execute a proportionate mitigation, rollback, repair or bounded fix. • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green. • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication. • Join customer conversations occasionally when direct technical involvement is genuinely useful. • Coordinate the right response • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix. • Escalate with evidence, customer impact, actions already taken and the specific decision or help required. • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait. • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners.
• Respond to live incidents • Maintain accurate incident records and handover • Improve the reliability system • Coordinate the right response • Protect developers from routine pages • Deliver timely and accurate impact assessments • Use AI throughout triage and diagnosis • Produce a clear incident record • Validate that customer outcome has recovered • Keep ownership, uncertainty, decisions and next actions visible
• Respond to live incidents • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state. • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code. • Use AI throughout triage and diagnosis while checking its conclusions against real evidence. • Choose and execute a proportionate mitigation, rollback, repair or bounded fix. • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green. • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication. • Join customer conversations occasionally when direct technical involvement is genuinely useful. • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix. • Escalate with evidence, customer impact, actions already taken and the specific decision or help required. • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait. • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners. • Remove, consolidate and tune low-value alerts, and design monitoring around real service and customer outcomes. • Analyse material incidents with AI, validate the conclusions and turn repeated failure patterns into better alerts, runbooks, AI Skills, automation or product improvements. • Improve dashboards, diagnostics, service ownership and escalation information so common incidents are easier to understand and resolve. • Create safe, supervised automation for common operational actions. • Work with product teams to close observability, rollback, runbook and supportability gaps.
Role Description Production Support (Spring Boot, NodeJS, Python) Location: Mexico, remote 100% - Excellent and strong communication skills specifically in debug production systems. - Minimum three years of experience building and supporting backend services with NodeJS. - Minimum three years of experience in Spring Boot 3.X. - Minimum three years of experience in Python for data processing, scripting, automation, and backend services. - Minimum three years working with relational or NoSQL databases and building scalable web applications. - Minimum five years of experience in SQL specifically PostgreSQL. - Knowledge of TypeScript and ability to write clean, maintainable, strongly typed code. - Solid frontend development experience using React and modern React patterns. - Strong hands-on experience with Material UI both current and older versions. - Experience designing and building REST APIs and integrating them with frontend applications. - Experience with testing frameworks for both frontend and backend development. Qualifications - Minimum three years of experience in Spring Boot 3.X. - Minimum three years of experience in Python for data processing, scripting, automation, and backend services. - Minimum three years of experience building and supporting backend services with NodeJS. - Minimum five years of experience in SQL specifically PostgreSQL. - Knowledge of TypeScript. - Solid frontend development experience using React. - Strong hands-on experience with Material UI. Requirements - Excellent communication skills. - Experience with relational or NoSQL databases. - Experience designing and building REST APIs. - Experience with testing frameworks for both frontend and backend development. Company Description
