SimScale empowers engineers with AI-native cloud simulation, enabling thousands of engineering decisions in seconds.
Senior SRE / Platform Engineer
Location
Germany
Posted
52 days ago
Salary
0
Seniority
Senior
Job Description
Senior SRE / Platform Engineer
SimScale
• Evolve our Kubernetes platform: Evaluate and adopt technologies such as Kubernetes Gateway API and service mesh patterns, and coordinate platform evolution across 10+ engineering teams. • Take observability to the next level: Drive organization-wide adoption of OpenTelemetry for distributed tracing and metrics, and help teams define meaningful SLOs. • Shape multi-region architecture and data residency: Support our move from an EU-centered footprint toward a global, multi-cloud architecture that satisfies disaster-recovery and data-residency requirements. • Own cloud cost and efficiency at scale: Keep petabyte-scale infrastructure cost-efficient, secure, and well-instrumented. • Improve tooling: Build self-service AWS account provisioning, guardrails and AI-assisted automations that help engineering teams manage infrastructure safely and efficiently at scale.
Job Requirements
- 5+ years of professional experience in SRE, platform, or infrastructure engineering.
- Software development experience: Your background is rooted in software development, and you moved into SRE from there. You write production-quality software in at least one of Python, Go, Rust, or Java.
- Strong systems foundation: You understand Linux internals and distributed systems well enough to debug complex production behavior.
- Hands-on cloud and infrastructure experience: AWS (or GCP), declarative infrastructure (Terraform), gitops-workflow (ArgoCD) and container orchestration (Kubernetes).
- Observability and reliability experience: You have worked with OpenTelemetry, Prometheus, distributed tracing, monitoring, and meaningful SLOs/SLIs.
- Production debugging depth: You can investigate complex failures, communicate clearly during incidents, and turn findings into durable improvements.
- Security and compliance awareness: You understand how infrastructure decisions affect access control, auditability, disaster recovery, logging, and standards such as SOC 2.
- Clear communication: You can explain trade-offs to engineering teams and help others adopt better platform practices without unnecessary friction.
Benefits
- Join a dedicated, supportive team with unlimited growth opportunities and leadership potential
- Make an impact quickly by sharing ideas and contributing to creative, goal-oriented projects
- Work in a diverse, inclusive environment with colleagues from over 35 countries
- Enjoy flexible hours and the freedom to work remotely from anywhere in the world
- Access comprehensive health coverage, retirement plans, paid time off, and wellness support
- Enjoy fresh office lunches or gift cards as a remote employee
- Grow as a professional with online/offline learning, language courses, and tech talks
- Connect at team events, join support groups, and contribute to our ESG and DE&I initiatives
- Participate in fun team challenges and competitions for added excitement and team spirit
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Senior DevOps
SoftExpert - Software for ExcellenceSoftware all-in-one para gestão da transformação digital, inovação e conformidade.
• Design, implement, and maintain AWS architectures focused on availability, scalability, and security. • Administer and optimize services such as EC2, EKS, RDS, S3, EFS, Lambda, CloudFront, and other platform resources. • Plan and execute migrations, upgrades, and infrastructure changes with minimal operational impact. • Manage complex network environments using VPCs, Transit Gateway, Site-to-Site VPN, Peering, and routing. • Configure and maintain Security Groups, NACLs, and access policies following security best practices. • Investigate and resolve connectivity, performance, and routing issues. • Implement and manage security and protection solutions in cloud environments. • Analyze and respond to security alerts, vulnerabilities, and risks. • Ensure infrastructure compliance through hardening, access management, and periodic reviews. • Administer Linux servers in production environments. • Perform patching, hardening, and lifecycle management of instances. • Create and maintain custom images for different use cases. • Develop and maintain infrastructure using Terraform. • Automate provisioning and configurations with Ansible. • Build and maintain CI/CD pipelines for infrastructure and applications. • Develop scripts to automate operational tasks. • Manage Kubernetes clusters (EKS), including upgrades, troubleshooting, security, and scalability. • Create and optimize Docker images. • Implement security policies and access control within clusters. • Monitor environments using observability tools. • Participate in incident management and production issue resolution. • Conduct root cause analyses (RCA) and implement preventive actions.
• Lead Cloud Migration: Spearhead the migration of workloads from AWS to Azure using an AKS-first (Azure Kubernetes Service) model. • Infrastructure as Code: Design and implement robust IaC for Azure environments. • Pipeline Standardization: Migrate, optimize, and standardize CI/CD pipelines. • Architecture Modernization: Refactor legacy infrastructure into cloud-native Azure architectures. • Risk Mitigation: Ensure minimal disruption during migration by designing efficient downtime and rollback plans. • Environment Optimization: Continuously optimize cloud environments for cost, performance, and scalability. • Cross-Functional Collaboration: Partner with engineering teams to modernize deployment workflows. • Documentation: Document migration patterns and create reusable templates for the broader team.
DevSecOps Engineer SR
Encora DigitalEncora, a leader in digital engineering, drives innovation by crafting cutting-edge, cloud-first, data-first, and AI-first solutions that redefine industries. S
Role Description We at Coforge are hiring a DevSecOps Engineer SR with the following skill set. - Design, implement, and maintain secure CI/CD pipelines for frontend, backend, and infrastructure delivery. - Automate infrastructure provisioning, configuration, and deployment processes for cloud and on-prem compatible environments. - Establish security controls across the software delivery lifecycle, including secrets management, dependency scanning, container security, and pipeline guardrails. - Implement secure perimeter and runtime patterns for the external portal, including API gateway controls, network segmentation, identity integration, and access policies. - Support deployment architectures for a scalable platform, secure external authentication, file distribution mechanisms, notifications, audit logging, and portal-to-enterprise platform integrations. - Define and maintain observability capabilities such as logs, metrics, alerting, and health monitoring for platform services and integrations. - Partner with engineering teams to embed security-by-design and operational best practices into application development. - Contribute to incident response readiness, vulnerability remediation workflows, compliance-aligned controls, and environment hardening. - Document deployment standards, platform architecture, runbooks, and operational procedures. Qualifications - Strong experience with DevOps and DevSecOps practices across build, release, deployment, and runtime operations. - Hands-on experience with CI/CD pipeline design, infrastructure as code, automation scripting, and environment provisioning. - Strong knowledge of cloud and hybrid deployment architectures, containerization, and orchestration concepts. - Solid understanding of application and infrastructure security, secrets management, IAM, vulnerability scanning, and secure delivery controls. - Experience implementing observability, alerting, logging, audit, and platform monitoring practices. - Familiarity with network security, gateway patterns, secure external access, and production hardening approaches. - Experience collaborating with software engineering teams to support secure and reliable software delivery. - Strong troubleshooting, documentation, and operational problem-solving skills. Requirements - Experience with Docker, Kubernetes, and secure container runtime practices. - Experience with cloud platforms such as AWS and secure storage/file distribution patterns. - Knowledge of SAST, DAST, dependency scanning, policy-as-code, and security compliance frameworks. - Experience supporting systems with multi-tenant isolation and external security requirements. - Familiarity with Python-based application environments, API-centric architectures, and enterprise integration patterns. Posted On 12-06-2026 At Coforge, we hire professionals based solely on their skills and do not discriminate based on age, disability, religion, gender, sexual orientation, socioeconomic status, or nationality.
• Own and continuously improve the infrastructure and operational systems that support Clever's engineering teams • Build safer deployment paths with automated health checks, documented rollback plans, and clear release visibility for high-risk services where the architecture supports them • Improve production reliability across infrastructure, data stores, queues, background jobs, logs, alerts, and supporting services • Lead or support infrastructure migrations, production cutovers, and environment changes across Heroku, AWS, ECS, RDS, networking, and DNS • Strengthen observability and incident response through better alerts, dashboards, escalation paths, post-incident follow-up, runbooks, and operational drills • Define practical reliability targets, service ownership expectations, and production-readiness standards for critical systems, internal tools, integrations, and AI-enabled workflows • Partner with engineers on infrastructure-sensitive application work, including scaling, error handling, deployment risk, data access, rollback planning, and small backend or internal-tooling changes • Support SOC2 and security readiness by improving infrastructure controls, access patterns, vulnerability handling, evidence collection, and repeatable operational practices • Improve local development and staging environments so engineers can onboard, test, and validate changes with less friction



