An ecosystem marketplace, subscription billing and management platform for distributors, vendors, MSPs and sellers
DevOps Engineer
Location
Latvia
Posted
3 days ago
Salary
0
Seniority
Senior
Job Description
DevOps Engineer
AppXite
• Operate and improve our Kubernetes-based infrastructure in close collaboration with software developers. • Own GitOps-based continuous delivery with ArgoCD and declarative, templated configuration using Kustomize, deployed through our CI/CD pipelines. • Manage Kubernetes clusters end-to-end, including provisioning, scaling, access control, and cost and performance optimization. • Run and maintain self-managed Kubernetes on-premises alongside Azure-managed clusters, including the networking and storage infrastructure that supports them. • Automate operational tasks and ensure our infrastructure remains reproducible, scalable, and well-documented. • Apply modern containerization and cloud-native best practices across the platform. • Troubleshoot and resolve operational issues in Kubernetes environments. • Document your work clearly and accurately.
Job Requirements
- 4+ years of experience in a DevOps role.
- Strong hands-on Kubernetes operations experience, including pods, services, deployments, ingress, persistent volumes, and maintaining healthy production clusters through upgrades, scaling, patching, and incident response.
- Experience with GitOps delivery using ArgoCD and declarative configuration with Kustomize, or strong equivalent experience that can be adapted quickly.
- Experience building and maintaining CI/CD pipelines; we use Azure DevOps.
- Solid experience with Docker and container registries.
- Comfortable running, or eager to grow into running, self-managed and on-premises Kubernetes environments, not only managed cloud clusters.
- Linux administration experience and scripting skills (Bash, Python, or similar).
- An automation-first mindset and a habit of solving problems by building scalable solutions.
- Ability to prioritize effectively, make sound decisions, execute independently, and collaborate as part of a team.
Benefits
- Remote-First Flexibility: Enjoy the freedom of a fully remote role, or choose to work from our Latvia office.
- Work with global Tech Leaders: Collaborate on solutions used by some of the biggest names in the industry, including Adobe, AWS, Cisco, Google, IBM, Microsoft, Lenovo, and Liquid.
- Skilled, International Team: Work alongside experienced professionals in a collaborative, multicultural environment. We value technical excellence, clear communication, and mutual support.
- Professional Development: Advance your skills with access to Microsoft certifications and continuous learning opportunities tailored to your role.
- Time to Recharge: Enjoy four weeks of paid vacation per year, plus paid public holidays.
- Referral Rewards: Help us grow our team - our employee referral program recognizes and rewards your contributions.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Role Description We run a distributed, camera-based video monitoring and AI alerting platform. The system spans the full spectrum of modern and legacy infrastructure: an AWS-hosted fleet of Java microservices and workers, a GPU-backed computer-vision inference pipeline, a real-time streaming and presence layer, and thousands of on-premise edge "media boxes" deployed in the field that ingest camera feeds, serve HLS video, and stream events back to the cloud. This is a reliability-first role. Your primary job is to keep a large, mixed operational estate healthy at scale: - Meaningful service objectives - Trustworthy alerting - Sound capacity - Tested disaster recovery - Fast, calm incident response If you think in SLOs, error budgets, and blameless postmortems, and you are happiest when a noisy, fragile system becomes quiet and predictable on your watch, this is your role. What you'll keep reliable - Computer-vision / AI models: frame-based inference services running on GPU EC2 (g4dn-class, AWS Deep Learning AMIs), fed by camera frames from S3 and a Kafka (Amazon MSK) event bus, with Redis (ElastiCache) for state. - Python services: the AI/alerts inference tier and supporting tooling. - Legacy Java services: ~140 Java 8 services and libraries (REST APIs, SQS/SNS workers, Lambda functions) running on Jetty 9.4, deployed to Elastic Beanstalk, ECS, and Lambda. - Edge appliances ("media boxes"): Ubuntu 22.04 / Docker Compose appliances managed remotely over AWS IoT Core secure tunnelling, with Cloudflare tunnels for egress. - Data and messaging backbone: PostgreSQL (RDS) across many schemas, TimescaleDB for analytics, Redis, DynamoDB, Amazon MSK (Kafka), SQS/SNS, and Kinesis. What you'll do Reliability and operations (the core of the role) - Own service objectives. Define SLIs and SLOs for the services that matter, manage error budgets, and use them to drive prioritization and change-rate decisions. - Make observability trustworthy. Own alert quality end-to-end; build the dashboards and the custom metrics, exporters, and instrumentation (CloudWatch, OpenTelemetry) needed to see the system clearly. - Lead incident response. Run incidents calmly, drive mean-time-to-recovery down, and produce blameless postmortems with action items that actually get closed. - Plan capacity and performance. Forecast and right-size compute (especially GPU), Kafka/MSK throughput and partitioning, RDS/TimescaleDB load, and Redis. - Own business continuity and disaster recovery. Backups, replication, failover, and recovery for RDS, MSK, Redis, and the edge fleet. - Keep the edge fleet healthy. Remote diagnosis and recovery over AWS IoT, container auto-update over systemd timers. - Engineer away toil. Write real software (Python, Golang, Bash) to automate operational work. - Govern production change safely. Enforce collaborative, reviewed change management. Delivery and platform (in support of reliability) - Keep CI/CD healthy and safe: CircleCI with Bazel/Gradle builds, OIDC-based AWS auth, container builds to ECR, and EB/ECS/Lambda deploys. - Maintain Terraform for the AWS estate (compute, networking, IAM, databases, messaging, monitoring). - Harden security and compliance: IAM least-privilege, Secrets Manager/KMS, TLS and certificate management. Qualifications - AWS certification is mandatory. A current AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect – Professional is strongly preferred. - 10+ years in Site Reliability Engineering or production operations at scale. - Demonstrated SLO/error-budget practice. - Strong production observability skills. - Proven incident command. - Capacity planning and performance experience across compute, databases, and a messaging or streaming system (Kafka/MSK ideal). - Disaster recovery ownership. - Software engineering ability for automation. - Expert with Terraform (or equivalent IaC) and strong Linux administration. - Database operations experience with PostgreSQL. - A reliability mindset. Requirements - Operating GPU workloads and serving computer-vision or ML models in production. - Apache MSK / Kafka and streaming-data operations. - AWS IoT Core at scale. - Managing a fleet of edge / on-premise devices. - Operating and modernizing legacy systems. - Chaos engineering / game-day practice, and capacity modeling. - Familiarity with Bazel in a monorepo; Cloudflare, Cognito/Auth0, API Gateway. How we work, and what we expect from this hire: - Observability must be trustworthy. - Change management is collaborative, never unilateral. - Follow-through over activity. - Prefer the right fix to a quick patch.
Site Reliability Engineer, US - Central/Eastern Timezone
PostHogProduct analytics, session replay, feature flags, A/B testing, data warehouse, CDP, surveys. PostHog does that.
• You won’t be in a typical “keep the lights on” SRE role. The work is about turning a fast-growing, stateful system into a predictable, well-automated platform. (provisioning, scaling, rebalancing, recovery) That means reducing operational stress, designing safe automation for traffic-heavy workloads, and building the tooling and patterns that let the system scale without scaling human effort. • You'll work on the kind of problems that only show up at large scale (petabytes of data, thousands of cores, constant ingestion) across a multi-region, multi-account AWS platform running many services on Kubernetes. • Operating EKS clusters across several environments with Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments • Managing and evolving a multi AWS account organization, provisioning, networking, access control, and cross-account connectivity • Maintaining the Terraform/Terragrunt IaC platform - modules, automated plan-on-PR / apply-on-merge pipelines, and safe patterns for shared infrastructure • Improving operational tooling around deploys, schema changes, backups, restores, and incident response • Reducing operational load by identifying repeat pain points and eliminating them through code and self-healing automation • Optimizing cloud spend as you go • Participating in on-call and incident response, with a strong focus on making incidents rarer over time. • You'll have room to design and automate, not just respond to alerts. You should join this team if you like deep ownership of production systems and enjoy building the platform layer that everything else runs on.
Site Reliability Engineer – ClickHouse
PostHogProduct analytics, session replay, feature flags, A/B testing, data warehouse, CDP, surveys. PostHog does that.
• Managing large fleets of EC2-based VMs, disks, and networking for data-intensive workloads • Improving operational tooling around deploys, schema changes, backups, restores, and incident response • Working closely with ClickHouse engineers to turn database-level needs into infra-level solutions • Reducing operational load by identifying repeat pain points and eliminating them through code and self-healing automation • Participating in on-call and incident response, with a strong focus on making incidents rarer over time • You’ll have room to design and automate, not just respond to alerts.
DevOps Engineer, Cloud Infrastructure and Live Games
Big Viking GamesBig Viking Games is a Canadian company based in London, Ontario with a second office location in Toronto, Ontario. Since 2011, Big Viking Games has been “maki
Title: DevOps Engineer, Cloud Infrastructure & Live Games Location: Toronto ON CA Job Description: Big Viking Games is a Canadian gaming company focused on building, operating, and growing long-standing online game communities. Our games have entertained players for years, supported by loyal audiences, live operations, evolving content systems, product innovation, and deep player-driven economies. Our flagship titles, YoWorld and FishWorld, have served millions of players over their lifetime. These are enduring live-service virtual worlds with rich in-game economies, virtual goods, social interaction, and long-term player engagement at their core. We are entering a new phase of modernization and growth, with a focus on stronger infrastructure, better automation, practical AI adoption, improved reliability, stronger security practices, and scalable systems that help our games and teams perform at a higher level. About the Role Big Viking Games is hiring a Senior DevOps Engineer to help design, maintain, secure, and modernize the infrastructure that supports our live-service games and internal development workflows. This is a hands-on role for someone who understands cloud infrastructure, automation, CI/CD, containers, monitoring, uptime, and production reliability — and who is comfortable working with legacy production systems alongside modern infrastructure patterns. Our games have been running for over a decade; the infrastructure reflects that history, and the right person sees that as an interesting challenge rather than a dealbreaker. You will work closely with engineering, product, QA, data, and live operations teams to improve how we build, deploy, monitor, and operate our systems. The right person is practical, security-aware, automation-minded, and able to balance speed with reliability. They can own infrastructure in a live production environment, improve DevOps processes, reduce manual work, and help development teams ship safely and efficiently. This is a hybrid role based in Toronto, with an expectation of working in office three days per week. Live-service games require operational awareness outside regular business hours, including periodic on-call and incident response availability. Requirements What You'll Do - Monitor, maintain, and improve cloud infrastructure across AWS, Netlify, Vercel, and related platforms that support our live games, data systems, internal tools, and AI-powered operational workflows. - Drive infrastructure modernization while maintaining uptime for live games with active player communities — every improvement ships while the plane is flying. - Build, maintain, and improve automation for deployments, environment management, provisioning, secrets rotation, and operational workflows — reducing manual toil and human error. - Implement and maintain Infrastructure as Code using tools such as Terraform, CloudFormation, CDK, or similar technologies. - Maintain and monitor data pipelines between game source databases (MariaDB), the Snowflake data warehouse, and downstream analytics and reporting systems — ensuring pipeline health, freshness, and alerting when data stops flowing. - Improve CI/CD pipelines, release workflows, and deployment reliability so development teams can ship safely and frequently. - Own secrets and credential lifecycle management across platforms — including API key rotation, access controls, environment variable governance, and least-privilege practices. - Support and improve the infrastructure that powers AI and automation tooling, including API integrations, MCP servers, serverless functions, webhook reliability, and orchestration platforms. - Improve observability across the stack: logging, metrics, alerting, dashboards, and operational visibility — with particular attention to early detection of silent failures in data pipelines and production systems. - Support incident response, root cause analysis, remediation planning, and post-incident improvements. - Help manage cloud spend, infrastructure usage, resource tagging, and environment efficiency. - Create clear documentation, runbooks, SOPs, and repeatable processes for infrastructure and DevOps workflows. What You Bring - 5+ years of experience in DevOps, infrastructure engineering, cloud engineering, site reliability engineering, or a similar role. - Strong hands-on experience with AWS or similar cloud platforms. - Experience designing, maintaining, and improving production infrastructure — including comfort with legacy systems that predate modern cloud-native patterns. - Proficiency with Infrastructure as Code tools such as Terraform, CloudFormation, CDK, Pulumi, or similar. - Experience with containerized applications, especially Docker. - Experience with CI/CD tools, version control, deployment automation, and modern release workflows. - Strong understanding of Linux systems, networking, cloud security, monitoring, logging, and operational troubleshooting. - Experience supporting production systems where uptime, reliability, and performance matter — especially systems that cannot tolerate extended downtime. - Experience with relational databases (MariaDB, MySQL, Postgres) and comfort working adjacent to data pipelines and ETL processes. - Security-aware mindset with practical experience in secrets management, credential rotation, access control, vulnerability reduction, and least-privilege practices. - Strong problem-solving skills and the ability to investigate complex infrastructure or production issues, including silent failures and data pipeline outages. - Ability to work closely with software engineers to improve build, deploy, and operational workflows. - Comfort creating documentation, runbooks, and repeatable operating processes. - Strong communication skills with both technical and non-technical stakeholders. - Practical ownership mindset with the ability to prioritize, execute, and close loops. - Nice to Have - Experience in gaming, live-service products, SaaS, digital products, or other high-availability consumer platforms. - Experience supporting live games, virtual worlds, multiplayer systems, or real-time online products. - Experience with GitHub Actions, GitLab CI, Jenkins, CircleCI, Buildkite, or similar CI/CD tools. - Experience with Datadog, Grafana, Prometheus, CloudWatch, ELK, OpenTelemetry, or similar observability tools. - Experience with Redis, Memcached, queues, workers, or event-driven systems. - Experience with Snowflake, data warehouse connectivity, ETL monitoring, or data pipeline reliability. - Experience with serverless platforms (Netlify Functions, Vercel, AWS Lambda) and multi-platform hosting environments. - Experience with disaster recovery, backup strategies, incident management, load testing, and performance tuning. - Experience improving cloud cost management, tagging, resource optimization, or infrastructure governance. - Experience with container orchestration platforms such as Kubernetes, ECS, EKS, or Nomad. - Experience operating infrastructure that supports AI/ML workflows, API integrations, or automation platforms (Make.com, webhook-driven orchestration, MCP servers). - Experience using AI tools such as Claude, ChatGPT, Gemini, or similar platforms to improve DevOps workflows, documentation, troubleshooting, and automation. - Experience working in small, high-leverage engineering teams where infrastructure ownership is broad and hands-on. - Ideal Candidate Profile - The ideal candidate is a practical infrastructure engineer who can keep live systems stable while helping modernize how the company builds, deploys, secures, and operates technology. - They are not only focused on tools. They understand uptime, developer experience, production risk, cloud costs, security, release quality, and operational discipline. They recognize that modernizing a decade-old live game requires patience, pragmatism, and the ability to improve systems incrementally without disrupting what's working. - They are comfortable operating across a mix of legacy and modern infrastructure, managing credentials and secrets lifecycle across multiple hosting platforms, and ensuring data pipelines are healthy and alerting properly. They can work independently, collaborate with engineers, and create systems that reduce friction instead of adding process for its own sake. - This role is best suited for someone who wants meaningful ownership over production infrastructure, cloud health, automation, and DevOps modernization inside a live-service gaming company. - Benefits Compensation Compensation range: $95,000 to $115,000 determined based on experience, technical depth, infrastructure ownership breadth, and overall fit. Benefits - Group Retirement Savings Plan matching and participation. - Comprehensive benefits package, including health, dental, and vision coverage. - Health and Wellness spending account. - Generous time off policies. - Opportunity to support long-running live-service games with established player communities. - Exposure to cloud modernization, DevOps automation, security improvement, and AI-enabled infrastructure workflows. - A high-impact role with meaningful ownership over reliability, performance, and engineering operations. Accessibility and Accommodation Big Viking Games is committed to creating an inclusive and accessible environment for all candidates. We welcome applications from individuals of all abilities and will provide accommodations throughout the hiring process as needed.


