The #1 Platform for Intelligent Banking Interactions
Senior Software Engineer – SRE, Observability Tooling
Location
Estonia
Posted
7 days ago
Salary
0
Seniority
Senior
Job Description
Senior Software Engineer – SRE, Observability Tooling
Glia
• Focus on building SRE and observability tooling — the platform, automation, and standards other teams use to keep their services healthy. • Developing standards, infrastructure and automation for dashboards, alerts, and monitors as code. • Partnering with development teams to establish production readiness and operational readiness. • Building the tooling and templates teams use to define, measure, and report on Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for their services. • Developing tooling to automate observability and operational workflows, eliminating manual toil for engineering teams. • Building and improving the incident response tooling and workflows that help teams resolve outages faster and learn from them.
Job Requirements
- Expert-level proficiency with AWS and Kubernetes (EKS), particularly in areas of observability, networking, and auto-scaling.
- Experience with modern observability platforms (e.g., DataDog, Prometheus) and a deep understanding of metrics, logging, and tracing.
- Deep, practical understanding of Site Reliability Engineering (SRE) principles (SLOs, error budgets, toil reduction).
- Demonstrable experience analyzing and troubleshooting large-scale distributed systems.
- Strong software development skills in a language like Python or Go, used to build operational tools, services, or automation.
- Expertise in designing and operating robust CI/CD pipelines for a microservices architecture (e.g., using ArgoCD, Github Actions, Helm).
- A systematic, data-driven approach to problem-solving and root cause analysis.
- Proficiency in using AI tools thoughtfully, maintaining ownership of the final output while recognizing the tools' limitations.
Benefits
- Flexible remote collaboration
- Optional offices in Tallinn and Tartu
- In-person innovation and connection sessions twice a year
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Senior Software Engineer, SRE / Observability Tooling
GliaThe #1 Platform for Intelligent Banking Interactions
• Build SRE and observability tooling • Develop standards, infrastructure, and automation for dashboards and alerts • Partner with development teams for production and operational readiness • Build tooling and templates for defining and reporting on Service Level Objectives (SLOs) • Automate observability and operational workflows • Improve incident response tooling and workflows
• Own and continuously improve the reliability of data pipelines across ingestion, transformation, and delivery layers, ensuring data is accurate, complete, and delivered on schedule. • Establish and maintain data reliability standards, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs) for both upstream ingestion and downstream data delivery. • Design, implement, and maintain comprehensive monitoring, logging, and observability frameworks for data pipelines, datasets, and data services with clear visibility into freshness, volume, schema changes, and data quality. • Design and implement data quality testing and validation frameworks — establishing test cases, golden datasets, and regression tests to detect quality issues early. • Lead incident response for data reliability issues, including detection, triage, communication, root cause analysis, and post-incident remediation with documented corrective actions. • Drive improvements in pipeline resiliency through retry strategies, backfills, idempotency, schema enforcement, and safe deployment practices.
Infrastructure / Systems Operations Engineer
ZignalyPassionate individuals at Zignaly are building a crypto investment platform to level the playing field for everyone.
• Operate and maintain our Linux, AWS, and blockchain-node infrastructure on a day-to-day basis. • Own routine administration, monitoring, and troubleshooting of services and networking. • Contribute to our infrastructure-as-code work with Terraform, growing your ownership over time. • Support blockchain node operations across our ecosystem. • Automate recurring tasks and diagnostics through scripting. • Help keep our systems secure, patched, and running smoothly.
Role Description As a Senior DevOps Engineer, you will work closely with Product, Engineering, and AI teams to shape our infrastructure strategy, design resilient cloud architectures, and ensure our platforms are secure, scalable, and high-performing. You will play a key role in bringing AI systems into production, enabling reliable delivery, strong observability, and operational excellence across our products and internal systems. - Design and operate secure, scalable, and high-quality infrastructure that supports modern applications and advanced AI workloads. - Build and maintain robust automation across CI/CD pipelines, infrastructure provisioning, and operational processes to improve reliability and minimize manual effort. - Integrate AI-driven solutions into operational workflows to enhance efficiency, detect anomalies, and accelerate delivery. - Apply strong systems engineering practices, including monitoring, incident management, performance optimization, and capacity planning. - Establish and uphold DevOps best practices, ensuring reproducibility, testing, documentation, and operational excellence. - Communicate technical decisions clearly and collaborate cross-functionally to support predictable delivery and effective problem-solving. - Provide mentorship and technical leadership, raising the level of platform engineering, DevOps maturity, and overall engineering quality across the organization. Qualifications - 6+ years of progressive experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Engineering. - Strong, hands-on experience across multi-cloud environments (AWS, GCP, Azure), including expertise in networking, compute, storage, security, and cost optimization. - Deep expertise in containerization and orchestration and extensive experience with Infrastructure as Code (IaC) (e.g., Terraform, Pulumi, CloudFormation). - Experience supporting or deploying AI/ML workloads (e.g., model inference, vector databases, GPU workloads), or strong familiarity with the infrastructure requirements for these systems. - Proven ability to design, build, and operate highly reliable, scalable production systems utilizing advanced Zero-Downtime Deployment Patterns (e.g., Blue/Green, Canary, progressive delivery, Preview Environments). - Expertise in modernizing deployments via GitOps practices (e.g., ArgoCD, Flux) and building Self-Service Developer Platforms that enable engineering efficiency (e.g., environment automation, internal tooling). - Experience implementing and managing Multi-Cloud API Gateways and Edge Routing solutions. - Strong background in platform security, including secrets management, Identity and Access Control (IAM), and Runtime/Security Hardening. - Solid understanding and practical experience with modern observability stacks. - Excellent communication and collaboration skills with a proven ability to describe complex infrastructure decisions clearly and a background in mentoring engineers and driving improvements in engineering practices. - Familiarity with modern programming languages like Node.js, NestJS, and Python is highly desirable for extending DevOps capabilities or integrating tooling.


