Storyteller logo
Storyteller

We built Stories so you don't have to.

Site Reliability Engineer

Production EngineerProduction EngineerContractRemoteSeniorTeam 11-50Since 2019H1B No SponsorCompany SiteLinkedIn

Location

Turkey

Posted

1 day ago

Salary

€32K / year

Seniority

Senior

Experience acceptedEnglish

Job Description

Site Reliability Engineer

Storyteller

• Respond to live incidents • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state. • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code. • Use AI throughout triage and diagnosis while checking its conclusions against real evidence. • Choose and execute a proportionate mitigation, rollback, repair or bounded fix. • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green. • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication. • Join customer conversations occasionally when direct technical involvement is genuinely useful. • Coordinate the right response • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix. • Escalate with evidence, customer impact, actions already taken and the specific decision or help required. • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait. • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners.

Job Requirements

  • Agency and ownership - You take responsibility for ambiguous live problems, gather evidence, choose a path and follow through after the immediate pressure has passed.
  • Operational judgement - You can separate customer impact, symptoms and likely causes, make practical decisions under uncertainty and recognise when an intervention is no longer safe or bounded.
  • Technical comfort and aptitude - You are comfortable exploring unfamiliar systems through code, logs, APIs, data, infrastructure and command-line tools, and can make hands-on changes with a clear validation plan.
  • AI-native execution - You use AI for substantive technical work - investigation, hypothesis generation, code, automation, incident analysis and workflow improvement - while supervising the agent and challenging its conclusions.
  • Accuracy and validation discipline - You actively look for false confidence and verify outcomes through appropriate technical and customer signals.
  • Systems thinking - You look for repeated patterns and improve the triggers, owners, runbooks, automation, metrics and feedback loops around the work.
  • Clear coordination and communication - You communicate calmly and concisely with Support, developers and non-technical stakeholders, making evidence, impact, uncertainty, ownership and next actions easy to understand.
  • Curiosity and resilience - You learn unfamiliar products and tools quickly, keep investigating when the first hypothesis fails and change your approach when the evidence demands it.
  • Previous responsibility for live production systems or an on-call rota is strongly preferred because it is useful evidence that you understand the realities of incident response. It is not an automatic requirement: we will also consider candidates who demonstrate exceptional ownership, judgement, technical aptitude, learning velocity and performance in the practical assessment.

Benefits

  • Fully remote working from anywhere in Turkey
  • Shared out-of-hours UK coverage, including active evening shifts and weekday overnight pager duty

Related Categories

Related Job Pages

More Production Engineer Jobs

Storyteller logo

Site Reliability Engineer

Storyteller

We built Stories so you don't have to.

Full TimeRemoteTeam 11-50Since 2019H1B No Sponsor

• Respond to live incidents • Maintain accurate incident records and handover • Improve the reliability system • Coordinate the right response • Protect developers from routine pages • Deliver timely and accurate impact assessments • Use AI throughout triage and diagnosis • Produce a clear incident record • Validate that customer outcome has recovered • Keep ownership, uncertainty, decisions and next actions visible

Egypt
€20K / year
Storyteller logo

Site Reliability Engineer

Storyteller

We built Stories so you don't have to.

ContractRemoteTeam 11-50Since 2019H1B No Sponsor

• Respond to live incidents • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state. • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code. • Use AI throughout triage and diagnosis while checking its conclusions against real evidence. • Choose and execute a proportionate mitigation, rollback, repair or bounded fix. • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green. • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication. • Join customer conversations occasionally when direct technical involvement is genuinely useful. • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix. • Escalate with evidence, customer impact, actions already taken and the specific decision or help required. • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait. • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners. • Remove, consolidate and tune low-value alerts, and design monitoring around real service and customer outcomes. • Analyse material incidents with AI, validate the conclusions and turn repeated failure patterns into better alerts, runbooks, AI Skills, automation or product improvements. • Improve dashboards, diagnostics, service ownership and escalation information so common incidents are easier to understand and resolve. • Create safe, supervised automation for common operational actions. • Work with product teams to close observability, rollback, runbook and supportability gaps.

Algeria
€20K / year

Role Description Production Support (Spring Boot, NodeJS, Python) Location: Mexico, remote 100% - Excellent and strong communication skills specifically in debug production systems. - Minimum three years of experience building and supporting backend services with NodeJS. - Minimum three years of experience in Spring Boot 3.X. - Minimum three years of experience in Python for data processing, scripting, automation, and backend services. - Minimum three years working with relational or NoSQL databases and building scalable web applications. - Minimum five years of experience in SQL specifically PostgreSQL. - Knowledge of TypeScript and ability to write clean, maintainable, strongly typed code. - Solid frontend development experience using React and modern React patterns. - Strong hands-on experience with Material UI both current and older versions. - Experience designing and building REST APIs and integrating them with frontend applications. - Experience with testing frameworks for both frontend and backend development. Qualifications - Minimum three years of experience in Spring Boot 3.X. - Minimum three years of experience in Python for data processing, scripting, automation, and backend services. - Minimum three years of experience building and supporting backend services with NodeJS. - Minimum five years of experience in SQL specifically PostgreSQL. - Knowledge of TypeScript. - Solid frontend development experience using React. - Strong hands-on experience with Material UI. Requirements - Excellent communication skills. - Experience with relational or NoSQL databases. - Experience designing and building REST APIs. - Experience with testing frameworks for both frontend and backend development. Company Description

Mexico
Full TimeRemoteTeam 11-50H1B Sponsor

Role Description As a Software Engineer specializing in block production, scheduling, and verification, you will play a critical role in fortifying Anza's Agave client and the broader Solana network. This team owns the core pipeline between networking and the runtime, responsible for processing ingested transactions and coordinating block-production and verification. Your work will directly contribute to the efficiency and reliability of our blockchain infrastructure, ensuring seamless and timely block production, scheduling, and verification. You will focus on optimizing the processes that underpin the generation and propagation of blocks, ensuring they are secure, performant, and scalable to meet the demands of future growth. Responsibilities - Develop and Optimize Block Production: Design, implement, and optimize the mechanisms for block production to enhance the throughput and stability of the Solana network. - Efficient Scheduling: Develop and refine scheduling mechanisms for both block production and block verification, enabling efficient transaction processing, fair resource allocation, and predictable performance. - Ensure Security and Integrity: Identify and mitigate potential security vulnerabilities within the pipeline, ensuring robust protection against emerging threats. - Scalability and Performance: Work on improving the scalability of the pipeline to handle increasing transaction volumes and validator participation without compromising on performance. - Testing and Validation: Create and execute comprehensive tests to validate the reliability and efficiency of the block production, block verification, and scheduling mechanisms, including stress tests, fault injection, and performance benchmarking. - Collaboration: Collaborate with cross-functional teams, including core protocol engineers, security experts, and infrastructure teams, to ensure the seamless integration and functioning of the block production components. - Documentation and Code Review: Maintain thorough documentation of the block production and scheduling protocols and conduct peer code reviews to uphold high standards of code quality and consistency. - Shared-Memory Systems: Design and maintain high-performance shared-memory interfaces between scheduling and execution stages, enabling low-latency coordination and efficient data movement across the validator pipeline. Qualifications - A Bachelor's degree in Computer Science, Engineering, or equivalent practical experience. - 3+ years of hands-on experience with core infrastructure software and distributed systems. - Strong proficiency in systems programming languages such as Rust, C, or C++. - Experience with distributed systems and blockchain technology is highly desirable. - Ability to analyze complex systems, identify potential issues, and develop effective solutions. - Knowledge of common security threats and best practices in securing block production processes. - Experience with performance profiling and optimization techniques. - Excellent teamwork and communication skills, with the ability to work effectively in a collaborative environment. Preferred Qualifications - Familiarity with Linux, systems automation tools, and systems architecture. - Deep understanding of architecture and principles underlying distributed systems. - A knack for designing secure protocols, software, and algorithms that minimize trust requirements. - Active participation in Bitcoin/Ethereum/Blockchain projects or the open-source community is highly desirable. Benefits - Dynamic, fast-paced environment focused on innovation and problem-solving. - Direct impact on the security and scalability of blockchain technology. - Contribution to the foundation of decentralized applications worldwide. - Competitive salary range for US-based candidates: $180,000 USD to $300,000 USD, determined throughout the interview process based on experience, skill, and location.

United Kingdom
$180K - $300K / year