Glia logo
Glia

The #1 Platform for Intelligent Banking Interactions

Senior Software Engineer, SRE / Observability Tooling

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 201-500Since 2012H1B No SponsorCompany SiteLinkedIn

Location

Poland

Posted

7 days ago

Salary

0

Seniority

Senior

Job Description

Senior Software Engineer, SRE / Observability Tooling

Glia

• Build SRE and observability tooling • Develop standards, infrastructure, and automation for dashboards and alerts • Partner with development teams for production and operational readiness • Build tooling and templates for defining and reporting on Service Level Objectives (SLOs) • Automate observability and operational workflows • Improve incident response tooling and workflows

Job Requirements

  • Expert-level proficiency with AWS and Kubernetes (EKS)
  • Experience with modern observability platforms (e.g., DataDog, Prometheus)
  • Deep understanding of Site Reliability Engineering (SRE) principles (SLOs, error budgets)
  • Experience analyzing and troubleshooting large-scale distributed systems
  • Strong software development skills (Python, Go)
  • Expertise in CI/CD pipelines (e.g., ArgoCD, Github Actions)
  • Data-driven problem-solving approach
  • Proficiency in using AI tools responsibly

Benefits

  • Health insurance
  • Retirement plans
  • Paid time off
  • Flexible work arrangements
  • Professional development

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Full TimeRemoteTeam 201-500H1B Sponsor

• Own and continuously improve the reliability of data pipelines across ingestion, transformation, and delivery layers, ensuring data is accurate, complete, and delivered on schedule. • Establish and maintain data reliability standards, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs) for both upstream ingestion and downstream data delivery. • Design, implement, and maintain comprehensive monitoring, logging, and observability frameworks for data pipelines, datasets, and data services with clear visibility into freshness, volume, schema changes, and data quality. • Design and implement data quality testing and validation frameworks — establishing test cases, golden datasets, and regression tests to detect quality issues early. • Lead incident response for data reliability issues, including detection, triage, communication, root cause analysis, and post-incident remediation with documented corrective actions. • Drive improvements in pipeline resiliency through retry strategies, backfills, idempotency, schema enforcement, and safe deployment practices.

United States
Zignaly logo

Infrastructure / Systems Operations Engineer

Zignaly

Passionate individuals at Zignaly are building a crypto investment platform to level the playing field for everyone.

DevOps Engineer7 days ago
Full TimeRemoteTeam 11-50Since 2018H1B No Sponsor

• Operate and maintain our Linux, AWS, and blockchain-node infrastructure on a day-to-day basis. • Own routine administration, monitoring, and troubleshooting of services and networking. • Contribute to our infrastructure-as-code work with Terraform, growing your ownership over time. • Support blockchain node operations across our ecosystem. • Automate recurring tasks and diagnostics through scripting. • Help keep our systems secure, patched, and running smoothly.

United Arab Emirates
$40K - $50K / year

Role Description As a Senior DevOps Engineer, you will work closely with Product, Engineering, and AI teams to shape our infrastructure strategy, design resilient cloud architectures, and ensure our platforms are secure, scalable, and high-performing. You will play a key role in bringing AI systems into production, enabling reliable delivery, strong observability, and operational excellence across our products and internal systems. - Design and operate secure, scalable, and high-quality infrastructure that supports modern applications and advanced AI workloads. - Build and maintain robust automation across CI/CD pipelines, infrastructure provisioning, and operational processes to improve reliability and minimize manual effort. - Integrate AI-driven solutions into operational workflows to enhance efficiency, detect anomalies, and accelerate delivery. - Apply strong systems engineering practices, including monitoring, incident management, performance optimization, and capacity planning. - Establish and uphold DevOps best practices, ensuring reproducibility, testing, documentation, and operational excellence. - Communicate technical decisions clearly and collaborate cross-functionally to support predictable delivery and effective problem-solving. - Provide mentorship and technical leadership, raising the level of platform engineering, DevOps maturity, and overall engineering quality across the organization. Qualifications - 6+ years of progressive experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Engineering. - Strong, hands-on experience across multi-cloud environments (AWS, GCP, Azure), including expertise in networking, compute, storage, security, and cost optimization. - Deep expertise in containerization and orchestration and extensive experience with Infrastructure as Code (IaC) (e.g., Terraform, Pulumi, CloudFormation). - Experience supporting or deploying AI/ML workloads (e.g., model inference, vector databases, GPU workloads), or strong familiarity with the infrastructure requirements for these systems. - Proven ability to design, build, and operate highly reliable, scalable production systems utilizing advanced Zero-Downtime Deployment Patterns (e.g., Blue/Green, Canary, progressive delivery, Preview Environments). - Expertise in modernizing deployments via GitOps practices (e.g., ArgoCD, Flux) and building Self-Service Developer Platforms that enable engineering efficiency (e.g., environment automation, internal tooling). - Experience implementing and managing Multi-Cloud API Gateways and Edge Routing solutions. - Strong background in platform security, including secrets management, Identity and Access Control (IAM), and Runtime/Security Hardening. - Solid understanding and practical experience with modern observability stacks. - Excellent communication and collaboration skills with a proven ability to describe complex infrastructure decisions clearly and a background in mentoring engineers and driving improvements in engineering practices. - Familiarity with modern programming languages like Node.js, NestJS, and Python is highly desirable for extending DevOps capabilities or integrating tooling.

Vietnam

Role Description Gestalte moderne Infrastruktur mit uns. Werde Teil der Contensi Software GmbH. Zur Verstärkung unseres Teams suchen wir eine:n Senior DevOps Engineer, der/die mit fundierter Erfahrung, Weitblick und Leidenschaft für Technologie den Unterschied macht. - Du arbeitest eng mit unseren Kunden aus Industrie, öffentlichem Sektor und Wissenschaft an anspruchsvollen Infrastrukturprojekten, von der Planung bis zum stabilen Betrieb. - Du übernimmst die Architektur, Einrichtung und Weiterentwicklung von Cloud- und OnPrem-Infrastrukturen, darunter Public-Cloud-Umgebungen wie AWS, Azure, GCP sowie Plattformen wie Kubernetes und OpenShift. - Du betreibst und wartest bestehende Systeme und Rechenzentrumsinfrastrukturen, inklusive Hybrid- und Multi-Cloud-Setups. - Du analysierst bestehende Architekturen und unterstützt bei der Migration und Modernisierung von Legacy-Umgebungen. - Du richtest GitOps- und CI/CD-Workflows ein (z. B. mit ArgoCD, GitLab CI, Azure DevOps) und begleitest unsere Kunden bei der Automatisierung ihrer Entwicklungs- und Betriebsprozesse. - Du berätst zu aktuellen Technologien und Best Practices, insbesondere im Bereich Containerisierung, Cloud-native Tools und Infrastructure-as-Code (Terraform, Ansible). - Du arbeitest konzeptionell mit an Themen wie Netzwerkdesign, Storage-Architekturen (z. B. Ceph, ODF, NFS) und Security-Standards (BSI, ISO27001, IAM, PKI). - Bei Wunsch betreust du Kunden auch vor Ort in Deutschland – und unterstützt unsere Kolleg:innen durch dein Fachwissen im Team. Qualifications - Mindestens 6 Jahre Berufserfahrung im Bereich DevOps, Infrastruktur oder Site Reliability Engineering. - Erfahrung mit Cloud-Plattformen: AWS, Azure, GCP. - Sehr gute Kenntnisse in Kubernetes und OpenShift. - Praktische Erfahrung mit einem oder mehreren CI/CD-Tools: Jenkins, GitLab CI, GitHub Actions, Azure DevOps o. ä. - Kenntnisse in Logging- und Monitoring-Systemen wie ELK, Datadog, Prometheus, New Relic o. ä. - Sehr gute Kenntnisse in Infrastructure as Code (Terraform, Ansible, Puppet). - Solides Netzwerkverständnis (TCP/IP, DNS, Routing, Firewalls etc.). - Erfahrung mit Storage-Systemen wie Ceph, MinIO etc. - Know-how in Security-Standards & Compliance (z. B. BSI, ISO27001, PKI, IAM). - Fundierte Kenntnisse in mindestens einer Programmiersprache: z. B. Java, Python, Go, Node.js oder TypeScript. - Du bist motiviert, dich kontinuierlich weiterzuentwickeln und neue Technologien zu erlernen. - Fliessende Deutsch- und Englischkenntnisse. Benefits - Spannende Kundenprojekte mit technologischer Tiefe. - Remote-Arbeit mit gelegentlichen vor Ort Besuchen beim Kunden. - Gestaltungsfreiraum - bring deine Ideen ein, übernimm Verantwortung und wachse mit uns. - Weiterbildungsbudget für Zertifikate, Konferenzen und Fachliteratur. - Modernstes Equipment (MacBook Pro & iPhone zur privaten/beruflichen Nutzung). - Zugang zu Lernplattformen wie Pluralsight und O’Reilly Online Learning. - Ergonomisches Homeoffice-Setup nach Wunsch. - Wellpass-Membership Zuschuss für sportliche Aktivitäten wie Schwimmbäder, Fitness-Studios uvm. - 30 Tage Urlaub.

Germany