Social Discovery Group (SDG) is one of the world's largest groups of social discovery companies, uniting millions of users across dozens of products. The compan
DevOps Engineer
Location
Worldwide
Posted
2 days ago
Salary
0
Seniority
Senior
Job Description
DevOps Engineer
Social Discovery Group - SDG
• Design and develop internal services and tools in Go • Design, build, and maintain scalable CI/CD pipelines using GitLab CI/CD • Improve and support Kubernetes-based environments • Contribute to the evolution of the internal deployment platform (Packman) • Optimize stage and ephemeral environments for performance, reliability, and cost-efficiency • Implement and maintain infrastructure as code using Terraform and Ansible • Improve observability through monitoring, logging, and alerting • Collaborate with engineering teams to enhance developer experience and delivery speed
Job Requirements
- Strong Go (Golang) skills with production experience
- 3–5+ years of experience in DevOps or Platform Engineering roles
- Strong hands-on experience with Kubernetes in production environments
- Solid experience with CI/CD tools (preferably GitLab CI/CD)
- Strong Linux administration skills
- Experience with Infrastructure as Code tools (Terraform, Ansible)
- Experience with observability tools (Grafana, ELK/OpenSearch, Prometheus/VictoriaMetrics)
- Good understanding of modern DevOps practices
- Fluent Russian and intermediate (B1) English level
Benefits
- REMOTE OPPORTUNITY to work full-time
- Vacation 28 calendar days per year
- 7 wellness days per year
- Bonuses up to $5000 for recommending successful applicants
- 50% payment for professional training
- Corporate discount for English lessons
- Health benefits
- Reimbursement of workplace costs up to $1000 gross every 3 years
- Internal gamified gratitude system
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Site Reliability Engineer - Monetization
XsollaXsolla's video game business engine helps game developers and publishers operate more efficiently and sell more games.
Title: Site Reliability Engineer (Monetization) Location: Montreal Type: Full time Workplace: remote Category: Infrastructure Team Job Description: ABOUT YOU We are looking for a Site Reliability Engineer (Monetization) who is pragmatic, product-minded, and equally comfortable writing code and running production systems to join our Infrastructure department's SRE team. The best candidate will be someone who thrives in a fast-paced, highly collaborative, and exceptionally dynamic setting and is excited to own the application-level infrastructure and reliability of a high-traffic commerce domain end to end - from deploy pipelines and Kubernetes manifests to SLOs, capacity planning, and production readiness. Strong Kubernetes, observability, and software engineering skills are essential, along with experience in operating production services in a cloud environment (GCP/GKE or comparable) and partnering closely with product development teams. The ability to hold a dual perspective - understanding both how developers ship features and what infrastructure needs to stay reliable - and to bring the reliability lens into design decisions early will be key to your success in this role. This is a hybrid embedded role: you remain part of the SRE organization (practices, standards, duty rotation) while being functionally embedded into the Monetization product domain. You'll build long-term working relationships with the domain's engineering teams, own a meaningful share of their application infrastructure execution, and co-author the reliability practices used company-wide. If you're passionate about making complex distributed systems boringly reliable and love building the commerce and monetization backbone that lets game developers around the world get paid, we would love to hear from you! ABOUT US Xsolla is a global commerce company with robust tools and services to help developers solve the inherent challenges of the video game industry. From indie to AAA, companies partner with Xsolla to help them fund, distribute, market, and monetize their games. Grounded in the belief in the future of video games, Xsolla is resolute in the mission to bring opportunities together, and continually make new resources available to creators. Headquartered and incorporated in Los Angeles, California, Xsolla operates as the merchant of record and has helped over 1,500+ game developers to reach more players and grow their businesses around the world. With more paths to profits and ways to win, developers have all the things needed to enjoy the game. Responsibilities - Own the application-level infrastructure of the Monetization domain: Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, and service-level networking and integrations - Own the domain's observability: design and implement SLOs/SLIs, monitors, alerts, and dashboards for critical services on Datadog and OpenTelemetry-based tooling - Help to set up and evolve CI/CD pipelines for domain services (GitLab CI, GitHub Actions), including deploy and rollback automation - Perform capacity planning and performance tuning ahead of expected load - product launches, sales events, and regional rollouts - including load testing and performance regression investigation - Run Production Readiness Reviews for new services and major changes; define and enforce what "production-ready" means for the domain - Support domain incident response: assist with deep investigation of complex incidents, contribute to post-mortems, drive follow-up reliability improvements, and maintain runbooks - Build domain-specific automation that reduces operational toil: runbook automation, deploy helpers, recurring operational scripts - Maintain and drive a forward-looking reliability roadmap for the domain together with product engineering leads - Participate in product team planning, refinements, and architecture reviews, bringing the reliability perspective before design decisions become expensive to change - Co-author company-wide SLO/SLI, capacity, and operational standards together with the broader SRE team; contribute improvements directly to shared SRE-operated subsystems - Participate in the SRE duty rotation, supporting developers across the company Qualifications & Skills - 3+ years of proven SRE, DevOps, or platform engineering experience: on-call or incident response duty, SLO/monitoring ownership, deploy pipeline and infrastructure work for production services - Software development background: you have built and shipped backend services, not only operated them - comfortable reading application code during an investigation and writing production-quality automation in at least one language (e.g., Go, PHP) - Hands-on Kubernetes experience: Helm, manifests, deploy strategies, debugging application-level performance and networking issues (GKE or another managed Kubernetes) - Solid observability practice: building monitors, dashboards, and SLOs/SLIs on a modern platform (Datadog preferred; Prometheus/Grafana experience also relevant), familiarity with OpenTelemetry - Infrastructure as Code exposure (Terraform/Terragrunt) for collaboration with platform teams - GCP experience (IAM, networking, managed services) - Experience building and maintaining CI/CD pipelines (GitLab CI and/or GitHub Actions) - Programming/scripting proficiency sufficient to build automation and tooling (e.g., Python, Go, or Bash) - Practical experience with incident response, post-mortems, and driving reliability improvements from incidents - Strong collaboration and communication skills — this role works embedded with product development teams daily - Experience in payments, fintech, e-commerce, or gaming — high-traffic transactional systems Nice to Have: - Kubernetes certifications - Google Cloud Platform certifications - HashiCorp certifications Salary varies depending on experience level and location. Benefits We are passionate about fostering a supportive environment for our team, so we prioritize the physical, mental, and emotional well-being of our employees and their families through a comprehensive Benefits Program. This includes medical, dental, and vision, PTO, and a personalized career roadmap for each employee. By investing in professional development through training and educational opportunities, we ensure that our team thrives both personally and professionally. Together, we're not just building a business; we're cultivating a community that values creativity, collaboration, and the transformative power of play. Equal Employment Opportunity Statement Xsolla is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, color, religion, sex, national origin, age, disability, sexual orientation, gender identity, or any other characteristic protected by law. We consider qualified applicants with criminal histories in accordance with the Fair Chance Act. Criminal History Consideration For the Site Reliability Engineer (Monetization) position, we will conduct a background check that may include the following: - Criminal history check - Employment verification - Education verification Relevance to Job Responsibilities The background check is relevant to this position because of the following role responsibilities: - Accessing confidential company data - Handling infrastructure that processes sensitive financial transactions - Ensuring compliance with regulatory requirements Rights Under the Fair Chance Act Applicants are encouraged to inquire about their rights under the Fair Chance Act. If you have questions regarding our hiring practices. By submitting the following job application form, you consent to Xsolla processing your data for career-related inquiries and potential employment opportunities. We process your data in accordance with this Xsolla Privacy Notice for Job Applicants.
• Diseño, implementación y mantenimiento de pipelines de CI/CD para automatizar los despliegues de aplicaciones. • Aprovisionamiento y administración de la infraestructura cloud principal (Azure) mediante Terraform. • Soporte, optimización y resolución de incidencias en entornos de Google Cloud Platform (GCP) cuando el proyecto lo requiera. • Monitoreo, escalabilidad y aseguramiento de la alta disponibilidad de clústeres de Kubernetes (AKS). • Coordinación directa con los equipos de desarrollo y arquitectura sobre aspectos técnicos y de despliegue. • Investigación y propuesta de mejoras en seguridad, rendimiento y optimización de costos en la nube.
• Diseño, implementación y mantenimiento de pipelines de CI/CD para automatizar los despliegues de aplicaciones. • Aprovisionamiento y administración de la infraestructura cloud principal (Azure) mediante Terraform. • Soporte, optimización y resolución de incidencias en entornos de Google Cloud Platform (GCP) cuando el proyecto lo requiera. • Monitoreo, escalabilidad y aseguramiento de la alta disponibilidad de clústeres de Kubernetes (AKS). • Coordinación directa con los equipos de desarrollo y arquitectura sobre aspectos técnicos y de despliegue. • Investigación y propuesta de mejoras en seguridad, rendimiento y optimización de costos en la nube.
DevOps AWS (Mid/Senior)
GFT Technologies SEProcuramos uma pessoa que: Goste de trabalhar em equipe e seja colaborativa em suas atribuições; Tenha coragem para se desafiar e ir além, abraçando novas oportunidades de crescimento; Transforme ideias em soluções criativas e busque qualidade em toda sua rotina; Tenha habilidades de resolução de problemas; Possua habilidade e se sinta confortável para trabalhar de forma independente e gerenciar o próprio tempo; Tenha interesse em lidar com situações adversas e inovadoras no âmbito tecnológico. Big enough to deliver – small enough to care. #VempraGFT #VamosVoarJuntos #ProudToBeGFT
Role Description Profissional de nível Pleno/Sênior que atue com (Devops/AWS). - Liderar a instrumentação ponta a ponta de métricas (infraestrutura e aplicação), logs estruturados e tracing distribuído, garantindo visibilidade do ecossistema. - Implementar, evoluir e gerenciar ferramentas de Application Performance Monitoring (APM) para identificar gargalos e otimizar a experiência do usuário. - Definir, implementar e monitorar SLIs, SLOs e Error Budgets, apoiando o equilíbrio entre inovação e estabilidade. - Planejar e executar experimentos de Chaos Engineering para validar hipóteses de falha e fortalecer a arquitetura. - Definir políticas de alertas preditivos, reduzindo fadiga de alertas e acelerando a resposta a incidentes. - Atuar de forma transversal apoiando pipelines de Engenharia de Dados e arquiteturas de microsserviços (APIs REST) na AWS. Qualifications - Domínio avançado de Datadog para dashboards, monitores, APM e Log Management. - Experiência sólida com AWS, Docker e Kubernetes (EKS). - Vivência em automação de infraestrutura e monitoramento. - Conhecimento prático dos pilares de SRE, incluindo gerenciamento de incidentes, SLIs, SLOs e Error Budgets. - Conhecimento de arquitetura de sistemas distribuídos, Alta Disponibilidade (HA), tolerância a falhas e resiliência. Requirements - Infraestrutura como Código com Terraform, Pulumi ou CloudFormation. - Observabilidade em Engenharia de Dados, incluindo pipelines de Big Data (Apache Airflow, Spark ou similares). - Conhecimento em SecOps/DevSecOps e Observability-driven Security. - Certificações AWS (ex.: AWS Certified DevOps Engineer) ou Kubernetes (CKA/CKAD). Benefits - Goste de trabalhar em equipe e seja colaborativa em suas atribuições. - Tenha coragem para se desafiar e ir além, abraçando novas oportunidades de crescimento. - Transforme ideias em soluções criativas e busque qualidade em toda sua rotina. - Tenha habilidades de resolução de problemas. - Possua habilidade e se sinta confortável para trabalhar de forma independente e gerenciar o próprio tempo. - Tenha interesse em lidar com situações adversas e inovadoras no âmbito tecnológico.

