Job Closed
This listing is no longer active.
The easier way to employ globally. Remote builds belonging for your team with payroll, benefits, & compliance solutions.
Staff Site Reliability Engineer I
Location
Worldwide
Posted
73 days ago
Salary
$188.6K - $212.2K / year
Seniority
Lead
No structured requirement data.
Job Description
Staff Site Reliability Engineer I
Remote
Role Description As a Staff SRE at Remote, you will own the technical direction of our SRE platform, shaping its architecture, reliability strategy, and long-term evolution. This is a leadership role as much as a technical one: - Drive platform-wide initiatives. - Set the reliability bar for engineering teams across the organization. - Be a force multiplier for the engineers around you. A key part of this role is identifying and leading opportunities to leverage AI: - Reduce operational toil. - Enable engineering teams to build, ship, and operate software more effectively. You will work with a high degree of autonomy, translating technical risks into business impact and aligning with Engineering Managers, Team Leads, and Product teams to ensure reliability and engineering efficiency are built into everything we do. Qualifications - 8+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering. - Deep expertise in Kubernetes: operating, designing, and scaling production clusters. - Proven experience designing and managing cloud infrastructure on AWS (or other cloud providers) at scale. - Strong infrastructure-as-code practice with Terraform. - Experience defining and operating reliability frameworks: SLOs, SLIs, error budgets, alerting strategies. - Solid observability background: Datadog, Grafana/Prometheus, or similar. - Proficiency with CI/CD platforms (GitLab CI, GitHub Actions, or similar) and deployment automation. - Comfortable with Bash and scripting for automation; broader programming skills are a plus. - Experience with container tooling (Docker) and the broader ecosystem around it. - Curiosity and practical experience applying AI tools to infrastructure, operations, or developer tooling. Requirements - Proven track record of driving platform-wide technical initiatives and influencing engineering direction without formal authority. - Strong communicator: able to tailor messaging to technical and non-technical audiences, write clearly, and align stakeholders across teams. - Self-directed: able to identify what needs attention, define the path forward, and execute with minimal supervision. - Experience mentoring senior engineers and creating space for others to lead and grow. - Comfortable navigating ambiguity, translating vague requirements into concrete solutions. - Approaches technical problems with a business lens, understands the cost and value of engineering decisions. Key Responsibilities - Own the technical direction of Remote's SRE/Platform domain, its architecture, tooling, and long-term roadmap. - Define and drive the reliability strategy across the platform: SLOs/SLIs, error budgets, observability, and incident management maturity. - Lead complex, cross-team infrastructure initiatives from discovery through delivery, delegating effectively and keeping projects aligned with business goals. - Identify and lead AI enablement initiatives across the engineering organization. - Drive AI-powered automation for platform operations: intelligent alerting, automated incident triage, self-healing infrastructure, and AI-assisted runbooks. - Contribute to capacity planning and cost-efficiency of Remote's infrastructure. - Mentor senior engineers, raising the technical bar through code reviews, design feedback, and hands-on guidance. - Collaborate with the Security team on platform hardening, threat mitigation, and compliance. - Be a steward of engineering quality across the SRE team, championing best practices, managing technical debt deliberately, and raising standards over time. - Contribute to hiring, onboarding, and continuously improving how the SRE team operates. Benefits - Work from anywhere. - Flexible paid time off. - Flexible working hours (we are async). - 16 weeks paid parental leave. - Mental health support services. - Stock options. - Learning budget. - Home office budget & IT equipment. - Budget for local in-person social events or co-working spaces.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
(Senior) Cloud Site Reliability Engineer (Platform)
Scalable GmbHScalable Capital is a leading digital investment and banking platform with a full banking licence, empowering people across Europe to shape their own finances. Scalable Broker makes it easy and affordable for clients to invest professionally in stocks, ETFs, cryptocurrencies, and derivatives, as well as set up savings plans. Scalable Wealth, the digital wealth management service, offers clients professional investment in ETF portfolios, and is also adopted as a white-label solution by banks and other B2B partners. The company’s offerings are rounded off by attractive interest rates, loans, and private equity. With the European Investor Exchange, Scalable Capital offers an exchange specifically for retail investors. Over one million clients have already entrusted more than €30 billion to the platform. Founded in 2014, Scalable Capital now employs over 700 people across Munich, Berlin, Vienna, Milan, and London. Together with the founding and management team, including Erik Podzuweit and Florian Prucker, they are working on a new generation of financial services.
Role Description Our team's mission is to provide secure, compliant and scalable building blocks or automations to enable developers to build workflows that rapidly but reliably ship software. Scalable Capital was built in the cloud from day one. Our services currently run on various AWS services like ECS, Fargate and Lambda and are distributed across multiple accounts. We embrace a DevOps culture where the development teams manage their CI/CD pipelines and cloud infrastructure for their services themselves. Our Platform Engineering Team focuses on providing everything necessary to build and ship code into these systems fast, security and developer friendly. - Shape the way how Scalable builds micro services in the most performant, secure and cost efficient way - Collaborate with cross-functional teams to identify and understand build and development requirements for our platform - Design and rollout CICD related improvements paired with internal automation tooling running in ECS, Lambda and EKS - Develop AI based supporting tools and libraries - Mentor and enable our software development teams to further foster our DevOps culture by educating them and providing reusable and unified building blocks which can be used to improve development speed, security, testing and releasing - Stay up-to-date with the latest industry trends, tools and techniques related to platform engineering - Design and implement best practices around building our infrastructure - Keep our internal development tooling up to date and support teams with migrating to new technologies Qualifications - Multiple years of experience with AWS and infrastructure as code (mainly Terraform) - Solid experience with GitHub Actions and Jenkins or similar CICD tools - Good working knowledge with Python and at least one additional general purpose programming language (Preferably Java/Kotlin or JavaScript/Node.js) and build automation tools - Solid understanding of scalable system design principles, distributed systems, and cloud technologies - A passion for automating, improving processes and working together with other developer teams - A degree in a relevant field of study (e.g. computer science, engineering, sciences) or work experience in a role that typically requires a university degree - Full professional proficiency in English and the ability to communicate concisely in an international English-speaking environment - Excellent communication and collaboration skills, with the ability to work effectively in a cross-functional team environment Benefits - Be part of one of the fastest-growing and most visible Fintech startups in Europe, creating innovative services that have a substantial impact on the lives of our customers - Work with an international, diverse, inclusive, and ever-growing team that loves creating the best products for our clients - Be productive with the latest hardware and tools - Learn and grow by joining our in-house knowledge sharing or career development sessions and spending your individual Education Budget - Learn and experience German culture first hand by joining our free German language classes - International relocation support is provided if required - Opportunity to work from abroad - Benefit from an attractive compensation package and from the company pension scheme - Monthly contribution of 50% for the ‘Deutschland Jobticket’ - Say goodbye to order commissions and say hello to your complimentary subscription of Scalable Capital's PRIME+ Broker - Enjoy flexible and discounted sports activities with Urban Sports Club
DevOps Engineer
LifeMDLifeMD (Nasdaq: LFMD) is a rapidly growing direct-to-consumer telemedicine company.
• Design, implement, and manage scalable, secure, and cost-effective cloud infrastructure primarily on AWS using Terraform • Develop and version control Terraform modules for automated provisioning, updating, and de-provisioning of cloud resources (e.g., EC2, S3, RDS, VPC, Lambda in AWS) • Design, build, and optimize automated CI/CD pipelines using GitHub Actions for various applications and microservices • Integrate automated testing, static code analysis, security scanning, and deployment steps into CI/CD workflows for high quality and secure releases • Implement, configure, and maintain comprehensive monitoring, logging, and alerting solutions (e.g., AWS CloudWatch, Datadog) for all environments • Develop custom dashboards, metrics, and alerts for real-time visibility into system health, performance, and security events • Proactively analyze logs and metrics to identify potential bottlenecks and issues • Participate in on-call rotations to swiftly respond to and resolve critical incidents, ensuring high service availability • Automate repetitive operational tasks, system configurations, and deployment processes using Python and Bash to enhance efficiency
Cloud Ops Engineer
SRM TechnologiesHelping automotive, healthcare, logistics & consumer sectors thrive with integrated Digital & Engineering solutions!
Role Description - Troubleshoot and resolve production and non-production issues within defined SLAs - Assess incident severity and respond appropriately to minimize customer impact - Analyze system logs and debug issues to restore services quickly - Perform root cause analysis (RCA) and recommend preventive measures - Communicate incident status and updates to stakeholders effectively - Deploy, monitor, and support application releases across environments - Identify recurring issues, analyze trends, and propose long-term solutions - Automate manual processes to improve operational efficiency - Participate in a 24x7 on-call rotation to ensure continuous support - Work effectively in a fast-paced, high-pressure environment with flexible schedules Qualifications - 5+ years of hands-on experience in application support or cloud operations - Strong knowledge of operating systems: Linux (OEL, CentOS, Ubuntu, Amazon Linux) and Windows - Experience with middleware technologies: JBoss, Apache, Tomcat, NGINX - Familiarity with DevOps tools: Jenkins, Terraform, Git, Ansible - Scripting skills in Bash/Shell, Python, Perl, or Groovy - Knowledge of Kubernetes (K8), ArgoCD, HelmChart - Knowledge & Experience on JIRA, Confluence - Solid understanding of networking concepts: DNS, firewalls, SFTP, file systems, LDAP - Experience with containerization technologies (e.g., Docker, Kubernetes) - Hands-on experience with monitoring tools such as Zabbix, Prometheus, Grafana, ELK stack - Experience working with AWS, OCI, or other cloud platforms - Familiarity with collaboration tools like JIRA and Confluence Requirements - AWS or other cloud certifications - Experience with Agile methodologies (Scrum/Kanban) - Knowledge of databases (Oracle or other RDBMS) and data migration tools - Exposure to big data technologies like Hadoop or Spark - Experience with WebSphere, WebLogic - Familiarity with Elasticsearch, Kibana, EMR - Experience supporting tools like Informatica or Looker
Role Description This is a remote position. We are looking for an experienced Senior DevOps Developer (SR1) to lead and enhance our cloud infrastructure, CI/CD pipelines, and deployment processes at Spark Eighteen. In this role, you will be responsible for architecting scalable, secure, and highly available systems while driving continuous improvements across infrastructure reliability, performance, and cost efficiency. - Take end-to-end ownership of deployment health, including SLOs and SLAs. - Work closely with Development and QA teams to streamline release automation. - Mentor junior engineers and contribute to architectural decision-making. - Lead complex initiatives with minimal supervision. - Propose innovative solutions aligned with business objectives. Qualifications - Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field. - 3–5 years of professional DevOps experience with demonstrated career progression. - Advanced Linux administration and shell scripting expertise. - Strong knowledge of Git workflows, including advanced branching and collaboration strategies. - Deep expertise in Kubernetes, including Helm, StatefulSets, Horizontal Pod Autoscalers, and Network Policies. - Advanced Terraform experience, including module development, remote backends, and workspace management. - Extensive hands-on experience with AWS services such as EC2, S3, IAM, VPC, and CloudWatch. - Advanced Docker usage with strong Kubernetes container optimization and deployment strategies. - Proven expertise in building and maintaining complex CI/CD pipelines using Jenkins and GitHub Actions. - Advanced secrets management experience using AWS SSM and HashiCorp Vault. - Strong experience in setting up logging and alerting systems using ELK Stack, Prometheus, and Alertmanager. - Advanced cloud security implementation experience, including IAM roles, Key Management Service (KMS), and Web Application Firewall (WAF). - Hands-on experience with GitOps practices using tools such as ArgoCD and Flux. - Strong performance tuning skills for infrastructure and containerized environments. - Advanced observability practices covering metrics, logs, and distributed tracing. - Ability to lead improvements to CI/CD systems and deployment pipelines. - Experience designing and implementing resilient, secure, and scalable infrastructure solutions. - Proven ability to proactively identify and resolve infrastructure bottlenecks and performance challenges. - Experience owning deployment health, including managing SLOs and SLAs. - Experience conducting infrastructure audits and driving cloud cost optimization initiatives. - Strong understanding of high-availability architectures and robust rollback strategies. - Proven ability to collaborate closely with Development and QA teams to improve release automation. - Demonstrated experience mentoring mid-level and junior DevOps engineers and promoting best practices. - Strong technical leadership skills with experience contributing to architectural decisions. - Ability to lead complex project components with minimal supervision. - Experience developing and executing risk mitigation strategies for infrastructure and deployment challenges. - Ability to propose innovative technical solutions aligned with business goals. - Excellent cross-functional communication skills with the ability to lead technical discussions. - Strong mentorship and coaching capabilities for junior and mid-level team members. - Strategic thinking skills with the ability to propose innovative and scalable solutions. - Excellent documentation and knowledge transfer skills. - Ability to align technical solutions with broader business strategy. - Proactive problem-solving mindset focused on continuous improvement. - Strong leadership skills to guide team performance and technical direction. - Effective collaboration skills across engineering, QA, and business teams. - Strategic approach to risk management and mitigation. - Experience with multi-cloud or hybrid-cloud environments is a plus. - Exposure to incident management processes and on-call responsibilities. - Advanced scripting skills in Groovy, Python, or Go for CI/CD automation. - Experience with infrastructure testing tools such as Terratest or Inspec. - Advanced cloud cost analysis and optimization experience. - Contributions to open-source projects or possession of advanced technical certifications. Benefits - Comprehensive insurance coverage that gives you peace of mind, so you can focus on doing your best work. - Flexible work arrangements designed to support sustained productivity, personal well-being, and work-life balance. - Continuous learning and accelerated skill development through hands-on projects and mentorship from experienced industry leaders. - Global client exposure across 20+ countries, offering real-world experience with diverse markets and business environments. - Opportunity to work on high-impact, large-scale projects that have collectively generated over $1B in measurable business value. - Competitive, market-aligned compensation packages that recognize performance, expertise, and long-term contribution. - Monthly demo days that celebrate innovation, showcase your work, and give you a real voice in what we build. - Annual recognition programs and performance-driven awards in a truly meritocratic environment. - Referral bonuses that reward you for helping grow a strong, like-minded team. - A strong problem-solving culture with opportunities to tackle meaningful, real-world challenges. - A positive, people-first workplace that supports happiness, balance, and long-term growth.



