Impact communities. For good.
AI DevOps Engineer – Global
Location
United States
Posted
6 days ago
Salary
0
Seniority
Senior
Job Description
AI DevOps Engineer – Global
MGT
• Build, maintain, and optimize CI/CD pipelines for AI and machine learning deployments. • Deploy and manage containerized AI workloads using Docker and Kubernetes. • Monitor production environments, model performance, infrastructure health, and system reliability. • Collaborate with AI engineers, data scientists, and solution architects to streamline deployment processes. • Implement Infrastructure-as-Code practices to improve scalability, consistency, and reproducibility. • Manage cloud infrastructure and platform services across AWS, Azure, and GCP environments. • Enforce security, compliance, and access control standards for AI systems. • Troubleshoot infrastructure and deployment issues while supporting incident response efforts. • Create and maintain operational documentation, deployment procedures, and technical runbooks. • Improve observability, monitoring, logging, and alerting frameworks for AI platforms.
Job Requirements
- Hands-on experience deploying and managing machine learning models in production environments
- Strong knowledge of containerization technologies and orchestration platforms
- Experience building and maintaining CI/CD pipelines
- Hands-on experience with Infrastructure-as-Code tools and cloud-native environments
- Familiarity with monitoring, logging, and observability solutions
- Strong understanding of security best practices for cloud and AI infrastructure
- Excellent written and verbal English communication skills
- Ability to work independently in a fully remote, U.S.-aligned environment
- Strong troubleshooting, problem-solving, and cross-functional collaboration skills
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field preferred, or equivalent professional experience.
- Three (3) or more years of experience in DevOps, MLOps, platform engineering, or related infrastructure roles.
- Experience working with government, education, or other regulated public sector organizations preferred.
- Familiarity with compliance frameworks such as FedRAMP, NIST, or similar regulatory standards preferred.
- Experience supporting LLM deployment pipelines, generative AI infrastructure, or AI platforms preferred.
- Experience with MLOps frameworks and model lifecycle management preferred.
- Cloud certifications, including AWS, Azure, or GCP, are a plus.
- Experience working in consulting or client-facing technical environments preferred.
Benefits
- Flexible paid time off
- 5% 401K matching program
- Equity opportunities
- Incentive and bonus programs
- Up to 16 weeks of paid parental leave
- Flexible spending accounts
- Full-health benefits with base employee coverage fully funded, comprising:
- Medical, dental, and vision coverage
- Life insurance
- Short and long-term disability coverage
- Income protection benefits
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Standardize and own the Tech Ops Incident Management Platform using Jira Service Management across our production fleets. • Automate incident resolution workflows with AI tooling, including runbook generation and assignment. • Design and implement proper documentation for all Tech Ops incident processes to ensure a clean handover. • Build out fleet management and task automation to support global 24/7 remote operations. • Own and customize the data-driven observability layer via Grafana across all internal tech teams. • Work closely with key leadership stakeholders (Head of Infrastructure & Cloud, VP Engineering) to independently drive architecture decisions.
Senior Site Reliability Engineer
ClickHouseClickHouse is an open-source, column-oriented OLAP database management system.
• Collaborate with various engineering teams in ClickHouse to design and implement scalable, secure, and highly available systems for ClickHouse. • Establish and manage service level objectives (SLOs) and service level agreements (SLAs) for ClickHouse Cloud. • Ensure all the infrastructure components in ClickHouse Cloud (including Dataplane, Control Plane and ClickHouse Core) have monitoring and alerting in place to ensure timely detection and resolution of incidents. • Enhance and refine incident response processes and post-mortem analysis for any outages in ClickHouse Cloud including working with the support team to communicate to the impacted customers. • Continuously improve the reliability and performance of our ClickHouse services. • Plan, enable, and drive Chaos initiatives across Engineering teams, based upon internal priorities. • Manage on-call processes to respond to performance and reliability issues, and establish best practices for coordinating escalation to resolve issues and minimize downtime.
Senior Site Reliability Engineer
ClickHouseClickHouse is an open-source, column-oriented OLAP database management system.
• Collaborate with various engineering teams in ClickHouse to design and implement scalable, secure, and highly available systems for ClickHouse. • Establish and manage service level objectives (SLOs) and service level agreements (SLAs) for ClickHouse Cloud. • Ensure all the infrastructure components in ClickHouse Cloud (including Dataplane, Control Plane and ClickHouse Core) have monitoring and alerting in place to ensure timely detection and resolution of incidents. • Enhance and refine incident response processes and post-mortem analysis for any outages in ClickHouse Cloud including working with the support team to communicate to the impacted customers. • Continuously improve the reliability and performance of our ClickHouse services. • Plan, enable, and drive Chaos initiatives across Engineering teams, based upon internal priorities. • Manage on-call processes to respond to performance and reliability issues, and establish best practices for coordinating escalation to resolve issues and minimize downtime.
Infrastructure – DevOps Lead
BrahmaThe only account you'll ever need to secure, transact, and explore onchain like never before.
• Lead, mentor, and grow a team of 7 DevOps and Infrastructure engineers. • Drive agile delivery, sprint planning, and backlog prioritisation to align infrastructure deliverables with AI research and product roadmaps. • Establish best practices for Reliability Engineering, Infrastructure-as-Code (IaC), continuous integration, and incident post-mortems. • Manage high-density GPU clusters across a multi-cloud ecosystem optimised for large custom AI model training and real-time inference workflows. • Oversee infrastructure consumption, track cloud/hardware costs, negotiate vendor terms, and optimise GPU utilisation. • Serve as the senior technical escalation point for complex infrastructure incidents and architecture decisions. • Standardise platform deployments using Infrastructure as Code and modern container orchestration. • Partner with security stakeholders to ensure our AI training environments meet industry security standards.



