MGT logo
MGT

Impact communities. For good.

AI DevOps Engineer – Global

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 501-1,000Since 1974H1B SponsorCompany SiteLinkedIn

Location

United States

Posted

6 days ago

Salary

0

Seniority

Senior

Job Description

AI DevOps Engineer – Global

MGT

• Build, maintain, and optimize CI/CD pipelines for AI and machine learning deployments. • Deploy and manage containerized AI workloads using Docker and Kubernetes. • Monitor production environments, model performance, infrastructure health, and system reliability. • Collaborate with AI engineers, data scientists, and solution architects to streamline deployment processes. • Implement Infrastructure-as-Code practices to improve scalability, consistency, and reproducibility. • Manage cloud infrastructure and platform services across AWS, Azure, and GCP environments. • Enforce security, compliance, and access control standards for AI systems. • Troubleshoot infrastructure and deployment issues while supporting incident response efforts. • Create and maintain operational documentation, deployment procedures, and technical runbooks. • Improve observability, monitoring, logging, and alerting frameworks for AI platforms.

Job Requirements

  • Hands-on experience deploying and managing machine learning models in production environments
  • Strong knowledge of containerization technologies and orchestration platforms
  • Experience building and maintaining CI/CD pipelines
  • Hands-on experience with Infrastructure-as-Code tools and cloud-native environments
  • Familiarity with monitoring, logging, and observability solutions
  • Strong understanding of security best practices for cloud and AI infrastructure
  • Excellent written and verbal English communication skills
  • Ability to work independently in a fully remote, U.S.-aligned environment
  • Strong troubleshooting, problem-solving, and cross-functional collaboration skills
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field preferred, or equivalent professional experience.
  • Three (3) or more years of experience in DevOps, MLOps, platform engineering, or related infrastructure roles.
  • Experience working with government, education, or other regulated public sector organizations preferred.
  • Familiarity with compliance frameworks such as FedRAMP, NIST, or similar regulatory standards preferred.
  • Experience supporting LLM deployment pipelines, generative AI infrastructure, or AI platforms preferred.
  • Experience with MLOps frameworks and model lifecycle management preferred.
  • Cloud certifications, including AWS, Azure, or GCP, are a plus.
  • Experience working in consulting or client-facing technical environments preferred.

Benefits

  • Flexible paid time off
  • 5% 401K matching program
  • Equity opportunities
  • Incentive and bonus programs
  • Up to 16 weeks of paid parental leave
  • Flexible spending accounts
  • Full-health benefits with base employee coverage fully funded, comprising:
  • Medical, dental, and vision coverage
  • Life insurance
  • Short and long-term disability coverage
  • Income protection benefits

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Full TimeRemoteTeam 51-200Since 2011H1B No Sponsor

• Standardize and own the Tech Ops Incident Management Platform using Jira Service Management across our production fleets. • Automate incident resolution workflows with AI tooling, including runbook generation and assignment. • Design and implement proper documentation for all Tech Ops incident processes to ensure a clean handover. • Build out fleet management and task automation to support global 24/7 remote operations. • Own and customize the data-driven observability layer via Grafana across all internal tech teams. • Work closely with key leadership stakeholders (Head of Infrastructure & Cloud, VP Engineering) to independently drive architecture decisions.

Poland
ClickHouse logo

Senior Site Reliability Engineer

ClickHouse

ClickHouse is an open-source, column-oriented OLAP database management system.

DevOps Engineer6 days ago
Full TimeRemoteTeam 51-200Since 2016H1B Sponsor

• Collaborate with various engineering teams in ClickHouse to design and implement scalable, secure, and highly available systems for ClickHouse. • Establish and manage service level objectives (SLOs) and service level agreements (SLAs) for ClickHouse Cloud. • Ensure all the infrastructure components in ClickHouse Cloud (including Dataplane, Control Plane and ClickHouse Core) have monitoring and alerting in place to ensure timely detection and resolution of incidents. • Enhance and refine incident response processes and post-mortem analysis for any outages in ClickHouse Cloud including working with the support team to communicate to the impacted customers. • Continuously improve the reliability and performance of our ClickHouse services. • Plan, enable, and drive Chaos initiatives across Engineering teams, based upon internal priorities. • Manage on-call processes to respond to performance and reliability issues, and establish best practices for coordinating escalation to resolve issues and minimize downtime.

Germany
ClickHouse logo

Senior Site Reliability Engineer

ClickHouse

ClickHouse is an open-source, column-oriented OLAP database management system.

DevOps Engineer6 days ago
Full TimeRemoteTeam 51-200Since 2016H1B Sponsor

• Collaborate with various engineering teams in ClickHouse to design and implement scalable, secure, and highly available systems for ClickHouse. • Establish and manage service level objectives (SLOs) and service level agreements (SLAs) for ClickHouse Cloud. • Ensure all the infrastructure components in ClickHouse Cloud (including Dataplane, Control Plane and ClickHouse Core) have monitoring and alerting in place to ensure timely detection and resolution of incidents. • Enhance and refine incident response processes and post-mortem analysis for any outages in ClickHouse Cloud including working with the support team to communicate to the impacted customers. • Continuously improve the reliability and performance of our ClickHouse services. • Plan, enable, and drive Chaos initiatives across Engineering teams, based upon internal priorities. • Manage on-call processes to respond to performance and reliability issues, and establish best practices for coordinating escalation to resolve issues and minimize downtime.

United Kingdom
Brahma logo

Infrastructure – DevOps Lead

Brahma

The only account you'll ever need to secure, transact, and explore onchain like never before.

DevOps Engineer6 days ago
Full TimeRemoteTeam 11-50Since 2022H1B No Sponsor

• Lead, mentor, and grow a team of 7 DevOps and Infrastructure engineers. • Drive agile delivery, sprint planning, and backlog prioritisation to align infrastructure deliverables with AI research and product roadmaps. • Establish best practices for Reliability Engineering, Infrastructure-as-Code (IaC), continuous integration, and incident post-mortems. • Manage high-density GPU clusters across a multi-cloud ecosystem optimised for large custom AI model training and real-time inference workflows. • Oversee infrastructure consumption, track cloud/hardware costs, negotiate vendor terms, and optimise GPU utilisation. • Serve as the senior technical escalation point for complex infrastructure incidents and architecture decisions. • Standardise platform deployments using Infrastructure as Code and modern container orchestration. • Partner with security stakeholders to ensure our AI training environments meet industry security standards.

United Kingdom