Storyteller logo
Storyteller

We built Stories so you don't have to.

Site Reliability Engineer

Production EngineerProduction EngineerContractRemoteSeniorTeam 11-50Since 2019H1B No SponsorCompany SiteLinkedIn

Location

Tunisia

Posted

2 days ago

Salary

€25K / year

Seniority

Senior

EnglishAzureCloud

Job Description

Site Reliability Engineer

Storyteller

• Respond to live incidents • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state. • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code. • Use AI throughout triage and diagnosis while checking its conclusions against real evidence. • Choose and execute a proportionate mitigation, rollback, repair or bounded fix. • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green. • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication. • Join customer conversations occasionally when direct technical involvement is genuinely useful. • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix. • Escalate with evidence, customer impact, actions already taken and the specific decision or help required. • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait. • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners. • Remove, consolidate and tune low-value alerts, and design monitoring around real service and customer outcomes. • Analyse material incidents with AI, validate the conclusions and turn repeated failure patterns into better alerts, runbooks, AI Skills, automation or product improvements. • Improve dashboards, diagnostics, service ownership and escalation information so common incidents are easier to understand and resolve. • Create safe, supervised automation for common operational actions. • Work with product teams to close observability, rollback, runbook and supportability gaps. • Detect and help contain unusual service-cost behaviour, then route wider follow-up to the appropriate cost or product owner. • Make reliability and on-call performance easier for the company to understand and improve over time.

Job Requirements

  • Previous responsibility for live production systems or an on-call rota is strongly preferred because it is useful evidence that you understand the realities of incident response.
  • Operational judgement - You can separate customer impact, symptoms and likely causes, make practical decisions under uncertainty and recognise when an intervention is no longer safe or bounded.
  • Technical comfort and aptitude - You are comfortable exploring unfamiliar systems through code, logs, APIs, data, infrastructure and command-line tools, and can make hands-on changes with a clear validation plan.
  • AI-native execution - You use AI for substantive technical work - investigation, hypothesis generation, code, automation, incident analysis and workflow improvement - while supervising the agent and challenging its conclusions.
  • Accuracy and validation discipline - You actively look for false confidence and verify outcomes through appropriate technical and customer signals.
  • Systems thinking - You look for repeated patterns and improve the triggers, owners, runbooks, automation, metrics and feedback loops around the work.
  • Clear coordination and communication - You communicate calmly and concisely with Support, developers and non-technical stakeholders, making evidence, impact, uncertainty, ownership and next actions easy to understand.
  • Curiosity and resilience - You learn unfamiliar products and tools quickly, keep investigating when the first hypothesis fails and change your approach when the evidence demands it.
  • Nice to have Cloud platforms such as Azure or Cloudflare.
  • Distributed application and API diagnostics.
  • Databases, queues and background-processing systems.
  • Observability, alerting and incident-management platforms.
  • Infrastructure, deployment and release automation.
  • Application development and safe production debugging.
  • AI coding agents and workflow automation.

Benefits

  • Fully remote working from anywhere in Tunisia!

Related Categories

Related Job Pages

More Production Engineer Jobs

Storyteller logo

Site Reliability Engineer

Storyteller

We built Stories so you don't have to.

ContractRemoteTeam 11-50Since 2019H1B No Sponsor

• Respond to live incidents • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state. • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code. • Use AI throughout triage and diagnosis while checking its conclusions against real evidence. • Choose and execute a proportionate mitigation, rollback, repair or bounded fix. • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green. • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication. • Join customer conversations occasionally when direct technical involvement is genuinely useful. • Coordinate the right response • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix. • Escalate with evidence, customer impact, actions already taken and the specific decision or help required. • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait. • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners.

Turkey
€32K / year
Storyteller logo

Site Reliability Engineer

Storyteller

We built Stories so you don't have to.

Full TimeRemoteTeam 11-50Since 2019H1B No Sponsor

• Respond to live incidents • Maintain accurate incident records and handover • Improve the reliability system • Coordinate the right response • Protect developers from routine pages • Deliver timely and accurate impact assessments • Use AI throughout triage and diagnosis • Produce a clear incident record • Validate that customer outcome has recovered • Keep ownership, uncertainty, decisions and next actions visible

Egypt
€20K / year
Storyteller logo

Site Reliability Engineer

Storyteller

We built Stories so you don't have to.

ContractRemoteTeam 11-50Since 2019H1B No Sponsor

• Respond to live incidents • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state. • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code. • Use AI throughout triage and diagnosis while checking its conclusions against real evidence. • Choose and execute a proportionate mitigation, rollback, repair or bounded fix. • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green. • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication. • Join customer conversations occasionally when direct technical involvement is genuinely useful. • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix. • Escalate with evidence, customer impact, actions already taken and the specific decision or help required. • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait. • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners. • Remove, consolidate and tune low-value alerts, and design monitoring around real service and customer outcomes. • Analyse material incidents with AI, validate the conclusions and turn repeated failure patterns into better alerts, runbooks, AI Skills, automation or product improvements. • Improve dashboards, diagnostics, service ownership and escalation information so common incidents are easier to understand and resolve. • Create safe, supervised automation for common operational actions. • Work with product teams to close observability, rollback, runbook and supportability gaps.

Algeria
€20K / year

Role Description Production Support (Spring Boot, NodeJS, Python) Location: Mexico, remote 100% - Excellent and strong communication skills specifically in debug production systems. - Minimum three years of experience building and supporting backend services with NodeJS. - Minimum three years of experience in Spring Boot 3.X. - Minimum three years of experience in Python for data processing, scripting, automation, and backend services. - Minimum three years working with relational or NoSQL databases and building scalable web applications. - Minimum five years of experience in SQL specifically PostgreSQL. - Knowledge of TypeScript and ability to write clean, maintainable, strongly typed code. - Solid frontend development experience using React and modern React patterns. - Strong hands-on experience with Material UI both current and older versions. - Experience designing and building REST APIs and integrating them with frontend applications. - Experience with testing frameworks for both frontend and backend development. Qualifications - Minimum three years of experience in Spring Boot 3.X. - Minimum three years of experience in Python for data processing, scripting, automation, and backend services. - Minimum three years of experience building and supporting backend services with NodeJS. - Minimum five years of experience in SQL specifically PostgreSQL. - Knowledge of TypeScript. - Solid frontend development experience using React. - Strong hands-on experience with Material UI. Requirements - Excellent communication skills. - Experience with relational or NoSQL databases. - Experience designing and building REST APIs. - Experience with testing frameworks for both frontend and backend development. Company Description

Mexico