MARGO logo
MARGO

DATA, AI & DIGITAL EXPERTS

Network Reliability Engineer

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 201-500Since 2005H1B No SponsorCompany SiteLinkedIn

Location

Poland

Posted

44 days ago

Salary

zł200 - zł250 / hour

Seniority

Senior

Job Description

Network Reliability Engineer

MARGO

• Build a large AI infrastructure with monitoring, diagnosis, and remediation of production incidents • Troubleshoot high-impact production issues in collaboration with other engineering teams • Participate in an on-call rotation to handle incidents and ensure service continuity • Implement and maintain observability solutions to monitor AI infrastructure and application health • Contribute to AI infrastructure lifecycle management across different environments and countries • Promote and apply best practices in terms of stability, resiliency, scalability, and security • Maintain clear technical documentation for tools and procedures • Contribute to system and tool evolution based on production feedback • Collaborate closely with development teams to ensure infrastructure readiness • Participate in team rituals and knowledge-sharing initiatives

Job Requirements

  • Experience with Go or Python
  • Strong scripting skills (Bash, Python)
  • Hands-on experience with Linux systems (Ubuntu/Debian)
  • Preferred hands-on experience with GPU & HPC infrastructure
  • Knowledge of networking (VLAN/LAN, TCP/IP, DNS, BGP, load-balancing, IPv6, etc.)
  • Familiarity with monitoring and logging tools (Prometheus, Grafana, Elastic, etc.)
  • Comfortable with Infrastructure-as-Code (Ansible, Salt, AWX, etc.)
  • Experience managing relational databases (MariaDB)
  • Understanding of CI/CD pipelines (GitLab)
  • Comfortable with English (written and spoken)

Related Categories

Related Job Pages

More DevOps Engineer Jobs

DevSecOps Engineer

Thinkahead Consultant Psychologist Pty Ltd

We get to the heart of the matter.....real people......real solutions

DevOps Engineer44 days ago
Full TimeRemoteTeam 1-10H1B No Sponsor

• Architect, deploy, and maintain complex AWS environments using Terraform. • Enforce least-privilege IAM policies, manage strict Security Group routing, and implement defense-in-depth security features. • Design and optimize GitHub Actions workflows for continuous integration and continuous deployment. • Build CloudWatch dashboards, configure metric filters, and set up automated alerting for operational and security events. • Establish and maintain SDLC standards.

India
Satori Analytics logo

DevOps Engineer, GCP

Satori Analytics

Clarity in decision making through Data and AI.

DevOps Engineer44 days ago
Full TimeRemoteTeam 51-200H1B No Sponsor

**What Your Day Might Look Like:** - **Cloud infrastructure as code**: Own and extend our Terraform estate across multiple GCP environments (base, core, obs, dev, test, prod), including GKE clusters, Cloud SQL (Postgres/MySQL), networking, buckets, and IAM. Drive the in-progress "Neo" platform rollout and the cutover/retirement of legacy infrastructure. - **Kubernetes & containers**: Manage workloads on GKE, maintain Dockerfiles and Helm-style application configs for ~10 backend services, and tune autoscaling, resource limits, and pod disruption budgets. - **Maintain and improve our GitHub Actions pipelines**: PR checks (Python/JS lint, type-check, tests), Terraform prechecks, image builds and pushes, auto-deploy, and DB-migration labelling/gating. Reduce build times and flakiness, and make deploys self-service for product teams. - **Data & messaging infrastructure**: Operate Postgres, Redis, and Celery-based async workers; manage Alembic migrations, queue health, and backpressure for long-running simulation jobs. - **Observability**: Own our monitoring stack — Grafana dashboards, ClickHouse, Langfuse (LLM tracing), and Celery queue metrics. Build alerting and SLOs so we catch issues before customers do. - **Security & secrets**: Manage secret distribution, least-privilege IAM, and remediation tracking. Partner with engineering on findings in our security assessment process. - **Cost & reliability**: Keep an eye on cloud and LLM-proxy (LiteLLM) spend, right-size resources, and improve resilience of the simulation and evaluation pipelines. **You'll work with:** - Cloud: Google Cloud Platform (GKE, Cloud SQL, GCS, IAM); some AWS / IBM footprint - IaC: Terraform (>= 1.14), multi-environment root modules - Containers/orchestration: Docker, docker compose (local), Kubernetes / GKE - CI/CD: GitHub Actions - Backend: Python 3.13+ (managed with uv), Celery, FastAPI-style HTTP APIs; Node/Express services - Data: PostgreSQL, MySQL, Redis, ClickHouse - Observability: Grafana, Langfuse, custom Celery metrics - LLM infra: LiteLLM proxy

Greece
Satori Analytics logo

DevOps Engineer

Satori Analytics

Clarity in decision making through Data and AI.

DevOps Engineer44 days ago
Full TimeRemoteTeam 51-200H1B No Sponsor

Role Description Are you passionate about AI? 🤖 At Satori Analytics, we aim to change the world one algorithm at a time by bringing clarity to global brands through Data & AI. From cloud-based ecosystems for fintech to predictive models for airlines, our cutting-edge solutions cover the entire data lifecycle—from ingestion to AI applications. As a fast-growing scale-up, our team of 100+ tech specialists—including Data Engineers, Data Scientists, and more—delivers innovative analytics solutions across industries like FMCG, retail, manufacturing and FSI. Join us as we lead the data revolution in South-Eastern Europe and beyond! Together with a partnering company, we're looking for a DevOps / Platform Engineer to own and evolve the infrastructure that keeps this platform reliable (AI agent evaluation platform), observable, secure, and fast to ship to. You'll work closely with backend, ML, and frontend engineers to make deploying and operating services boring, repeatable, and safe. What Your Day Might Look Like: - Cloud infrastructure as code: Own and extend our Terraform estate across multiple GCP environments (base, core, obs, dev, test, prod), including GKE clusters, Cloud SQL (Postgres/MySQL), networking, buckets, and IAM. Drive the in-progress "Neo" platform rollout and the cutover/retirement of legacy infrastructure. - Kubernetes & containers: Manage workloads on GKE, maintain Dockerfiles and Helm-style application configs for ~10 backend services, and tune autoscaling, resource limits, and pod disruption budgets. - Maintain and improve our GitHub Actions pipelines: PR checks (Python/JS lint, type-check, tests), Terraform prechecks, image builds and pushes, auto-deploy, and DB-migration labelling/gating. Reduce build times and flakiness, and make deploys self-service for product teams. - Data & messaging infrastructure: Operate Postgres, Redis, and Celery-based async workers; manage Alembic migrations, queue health, and backpressure for long-running simulation jobs. - Observability: Own our monitoring stack — Grafana dashboards, ClickHouse, Langfuse (LLM tracing), and Celery queue metrics. Build alerting and SLOs so we catch issues before customers do. - Security & secrets: Manage secret distribution, least-privilege IAM, and remediation tracking. Partner with engineering on findings in our security assessment process. - Cost & reliability: Keep an eye on cloud and LLM-proxy (LiteLLM) spend, right-size resources, and improve resilience of the simulation and evaluation pipelines. Qualifications - 3+ years in DevOps / SRE / Platform Engineering, or strong backend experience with heavy infra ownership. - Solid hands-on Terraform (modules, state, multi-environment) and cloud experience (GCP preferred; AWS/Azure transferable). - Production Kubernetes experience: deployments, services, autoscaling, debugging pods, rollouts/rollbacks. - Strong Docker fundamentals and comfort writing/optimising Dockerfiles. - CI/CD pipeline design and maintenance (GitHub Actions, or equivalent like GitLab CI / CircleCI). - Comfortable scripting and reading code in Python and/or Bash; able to navigate a polyglot monorepo. - Operational experience with relational databases and managed database services (migrations, backups, performance). - A reliability mindset: monitoring, alerting, incident response, and writing runbooks. Requirements - Experience operating Celery / distributed task queues and Redis at scale. - Familiarity with LLM/AI infrastructure (model proxies, GPU scheduling, token/cost management). - Observability tooling depth (Grafana, Prometheus, ClickHouse, OpenTelemetry, Langfuse or similar tracing). - Security/compliance experience (IAM hardening, secret management, vulnerability remediation). - Cost-optimisation experience for cloud + third-party API spend. - Experience supporting a monorepo with multiple language ecosystems and editable/internal package dependencies. Benefits - Competitive salary. - Training budget to level up your skills from top tech partners like Microsoft, AWS, Salesforce, and Databricks – whether it’s certifications or courses, we’ve got you covered. - Private insurance, top-tier tech gear, and the chance to work with a stellar crew. Company Description Ready to create some data magic with us? Hit that apply button and let’s get started. ✨

Greece
Dailymotion logo

Senior DevOps Engineer

Dailymotion

The home for videos that matter

DevOps Engineer44 days ago
Full TimeRemoteTeam 201-500Since 2005H1B Sponsor

Role Description We're hiring our first dedicated Devops / MLOps Engineer. Our ML team runs +100 active AI projects on GCP and needs the industrialized infrastructure to match. ML team is moving fast and they need a dedicated expert to connect the dots, align on the right practices, and help both sides deliver at their best. You'll be that person. You'll work at the heart of an AI-first environment, where shipping models fast and reliably is the core business, not a side project. - Empower ML Engineers with the tools, infrastructure, and frameworks they need to iterate fast autonomously. - Accelerate time-to-market for production-ready ML products: seamless integration, proper service connections, access to data and resources. - Own ML CI/CD in close collaboration with the ML team, adapting existing frameworks to ML-specific needs, not just consuming them. - Keep ML Engineers in control of their models in production: monitor, troubleshoot, iterate, refine directly in prod, no staging/prod mirror. - Enable large-scale ML experimentation: robust, reproducible, scalable environments for both internal tests and A/B testing in production. - Deliver concrete MLOps building blocks (MLflow, Kubeflow, KubeRay...) and manage GPU infrastructure dynamically, you've seen L40 shortages mid-training, you know how to handle them. - Tackle technical debt on existing projects while laying the right foundations for what's next. - Be the technical mediator between ML and Backbone teams: understand both sides, propose solutions that stick. - Handle run responsibilities: on-call, post-mortems, level-1 failure analysis. Qualifications - Solid MLOps or DevOps background; projects shipped in prod matter more than years on a resume. - GCP expert: Vertex AI, GKE, GCS, BigQuery. - Full GitOps: FluxCD first, ArgoCD accepted. - Kubernetes in prod, not just in a lab. - Hands-on with real MLOps tools: MLflow, Kubeflow, KubeRay. - GPU-aware: you've managed GPU scarcity at scale during mass training runs. - Python is a must, Bash expected, Go or Rust a plus. - IaC (Terraform), containerization (Docker, Helm), observability (Prometheus, Datadog, Looker). - AI Coding Assistants (Claude, Cursor, Dust). - Data lifecycle management (cost, security, encryption). - Fluent in French and English. Requirements - Jupyter Notebooks, broader ML/AI ecosystem. - Data pipelines (Airflow, Dataflow, Kestra). - Redis clusters and infrastructure performance optimization. Benefits - Don't check every box? Apply anyway. - We're looking for the right person, not the perfect resume. If this role excites you, let's talk. - Dailymotion is an equal opportunity employer. All positions are open to people with disabilities. Need accommodations? Just let us know. Interview Process - HR Interview with Marvin (30mn). - Manager interview with Cédric & Thierry (1h). - Technical Interview with 2 Architects (1h): no coding test, just a real technical conversation. - ML Interview with Brice & Samuel (1h). - Final Interview with Alan, VP Platform (1h). Additional Information - Type of Contract: Permanent. - Department: Operations. - Compensation: EUR 75000 - EUR 85000 yearly.

France
€75K - €85K / year