Job Closed
This listing is no longer active.
Simplicity, Scalability and Security
Site Reliability Engineer – System & Network
Location
Switzerland
Posted
50 days ago
Salary
0
Seniority
Senior
Job Description
Site Reliability Engineer – System & Network
Exoscale
• Design and maintain Exoscale’s base operating systems and hypervisor fleets including their network stacks and security filtering layers. • Contribute to the routing & security automation systems implementations. • Help shape our image strategy, from bare metal to Virtual Machines. • Maintain the system container runtimes. • Improve our provisioning and deployment systems. • Help improve bare metal and hypervisor systems performance. • Contribute to the overall design and the architecture of the Exoscale platform systems. • Contribute to internal tooling development. • Improve our systems and processes to be scalable and highly available, helping achieve outstanding SLAs. • Participate in code & changes reviews. • Take part in the on-call roll after a training period.
Job Requirements
- Have a solid experience with Linux, its kernel and Systems.
- Are familiar with KVM virtualization.
- Are familiar with network routing and transport / encapsulation protocols: BGP, VXLAN and EVPN.
- Are at ease with routing daemons like FRR and Bird.
- Have a good knowledge of Linux filtering with iptables and nftables.
- Have a good knowledge of container runtimes.
- Have a good experience with Golang, Python.
- Are familiar with server hardware.
- Have experience with configuration management solutions and large scale infrastructure.
- Love to automate anything that could be.
- Are curious, autonomous and embrace learning new things every day.
- Are team players and are comfortable working in a distributed team.
- Have good English communication skills, written and spoken.
Benefits
- Flexible working hours and working from home.
- Autonomous working conditions with a lot of freedom to create.
- Modern working atmosphere and centrally located office with great public transport connection.
- Team events as well as training and further education.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Cloud Operations Engineer
Made4netCloud based Warehouse Management & Supply Chain Execution solutions
• Provide production support as part of a global follow-the-sun operations team - monitor systems, triage alerts, respond to incidents, and ensure service continuity across all environments. • Monitor infrastructure and application health across security and operations dashboards; detect, triage, and respond to performance, availability, and security alerts; produce and maintain SLA reporting. • Deploy, manage, and maintain AWS infrastructure with a primary focus on EC2-based workloads across Windows and Linux environments. • Configure and manage load balancers (ALB/NLB), including URL rewrite rules, routing policies, and SSL termination for web applications. • Design and maintain high-availability architectures on AWS, ensuring redundancy, multi-AZ deployments, and tested failover procedures. • Own backup and disaster recovery operations - scheduling, retention policies, monitoring, and regular restoration testing. • Automate configuration management and application update deployments using Ansible, ensuring consistency and minimal downtime across all environments. • Manage database infrastructure across environments - provisioning, patching, performance tuning, and backup/recovery operations. • Troubleshoot and configure networking components including VPC, subnets, routing tables, DNS, security groups, and firewall rules. • Handle incident escalation, maintain shift handover documentation, and contribute to post-mortems and continuous improvement of support processes. • Maintain infrastructure runbooks, operational documentation, and contribute to automation and tooling improvements. • Support the company’s ongoing transition toward microservices and containerized architectures, contributing operational knowledge and helping ensure smooth adoption in production environments.
• Serve as the primary technical owner for production reliability across U.S. customer environments. • Investigate and resolve complex issues spanning web applications, APIs, backend services, data pipelines, cloud infrastructure, and customer integrations. • Lead production incident response efforts, coordinating cross-functional teams to restore service and minimize customer impact. • Perform root cause analysis and drive corrective actions that improve long-term system stability and resilience. • Partner with software engineering and platform teams to identify recurring reliability risks and implement sustainable solutions. • Design, configure, and validate secure customer connectivity solutions including Site-to-Site VPNs, Transit Gateway integrations, routing configurations, and secure network paths. • Support customer onboarding initiatives by troubleshooting connectivity challenges and ensuring consistent implementation processes. • Enhance platform observability through improvements in monitoring, logging, alerting, tracing, and operational dashboards. • Contribute to CI/CD, infrastructure automation, and deployment processes that improve release safety and operational consistency. • Develop operational tooling that supports incident response, troubleshooting, onboarding, and system monitoring activities. • Collaborate with engineering leadership to improve cloud architecture, scalability, security, and operational readiness. • Partner with customer-facing teams to communicate technical issues, remediation plans, and reliability improvements in a clear and effective manner. • Support compliance, security, and risk management initiatives within highly regulated healthcare environments.
Estágio DevOps
Viasoft Korp | Industry ERPO sistema de gestão nascido na indústria que vive e respira processos industriais e distribuição 💙
• Auxiliar em rotinas envolvendo: - microsserviços; - containers; - observabilidade; - alta disponibilidade; - integração contínua; - entrega contínua; - monitoramento; - deploy contínuo. • Auxiliar em rotinas de automação de infraestrutura em ambientes cloud e on-premise; • Auxiliar na implantação e evolução de ambientes Kubernetes; • Auxiliar na automatização de processos utilizando Ansible; • Auxiliar na implementação, evolução de monitoramentos e observabilidade com Grafana; • Atuar no auxílio de resolução de incidentes e troubleshooting entre serviços e ambientes; • Auxiliar na garantia de estabilidade, disponibilidade e performance dos ambientes; • Auxiliar na evolução de pipelines e ferramentas de CI/CD; • Documentar procedimentos, fluxos e configurações; • Participar ativamente da evolução tecnológica da plataforma da Korp.
• Owning the SRE infrastructure lifecycle from design reviews and pre-rollout readiness assessments through production sign-off and ongoing reliability management • Designing and implementing frameworks that reflect customer experience for load balancing services and driving action when error budgets are at risk • Building and maintaining observability pipelines from load-balancing components and system-level sources to dashboards that enable rapid incident triage • Leading technical incident response for complex NB/NLB failures, acting as the technical commander and driving root cause analysis and preventive follow-through • Developing and automating safe deployment workflows for phased releases, including bake-period monitoring, feature flag management, and validation across global datacenter rollouts • Reviewing design documents, product-requirement documents and producing actionable SRE input on operational risks, capacity implications, Day-2 concerns, and product strategy gaps • Building automation and tooling using Python or Go that reduces operational toil and improves team-wide operational capability




