Strategic open source infrastructure for containers and virtual machines.
AI Infrastructure & Platform Operations Engineer
Location
California
Posted
7 days ago
Salary
0
Seniority
Senior
Job Description
AI Infrastructure & Platform Operations Engineer
Mirantis
• Monitor, operate, and support production AI infrastructure platforms • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents • Support NVIDIA GPU infrastructure and associated platform services • Monitor and troubleshoot Kubernetes-based environments • Investigate performance, availability, and reliability issues across infrastructure and platform components • Collaborate with engineering teams, hardware vendors, Data Center personnel, and service delivery teams to resolve technical issues • Participate in incident response, root cause analysis, and operational improvement activities • Contribute to improvements in monitoring, observability, automation, and operational processes • Maintain operational documentation, runbooks, and knowledge articles
Job Requirements
- 3+ years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or related technical roles
- Strong Linux administration and troubleshooting skills
- Good understanding of networking concepts and experience diagnosing infrastructure-related issues
- Working knowledge of Kubernetes in production environments
- Experience supporting production infrastructure and services
- Strong analytical and problem-solving skills
- Experience working within structured operational and incident management processes
- Excellent communication and collaboration skills
- Ability to work within a shift-based operational environment
Benefits
- Work with an established Silicon Valley leader in the cloud infrastructure industry
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies
- Be a part of cutting-edge, open-source innovation
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued
- Professional development and training
- Attend conferences and working groups
- Company outings, happy hours, hackathons, and tech talks
- Receive a competitive compensation package with a strong benefits plan
Related Guides
Related Categories
Related Job Pages
More Infrastructure Engineer Jobs
Azure Infrastructure Engineer
Bright Vision TechnologiesBright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.
Role Description We are seeking a skilled Azure Infrastructure Engineer to design, deploy, and operate large-scale, secure, and resilient cloud platforms on Microsoft Azure. In this role you will own the end-to-end cloud engineering lifecycle, including: - Architecture - Infrastructure-as-code - Automation - Security hardening - Cost optimization - Observability - Ongoing operational excellence for production workloads The ideal candidate will combine deep technical expertise across Azure services with strong DevOps engineering practices, and will partner closely with application development, security, and SRE teams to deliver cloud-native solutions that meet demanding business requirements for scalability, reliability, and compliance. Qualifications - Bachelor’s degree in Computer Science, Engineering, or a related technical discipline. - Five or more years of cloud engineering experience, with at least three years focused on Microsoft Azure in production environments. - Strong hands-on experience with Azure core services, including compute, storage, networking, identity, and platform-as-a-service offerings. - Production-level experience with infrastructure-as-code tools such as Terraform, Bicep, or ARM templates. - Solid experience designing and operating Azure Kubernetes Service (AKS) clusters at scale. - Hands-on experience with Azure DevOps or GitHub Actions for CI/CD across infrastructure and applications. - Strong scripting skills in PowerShell, Bash, and Python, with the ability to write maintainable automation code. - Deep understanding of cloud security principles, identity management, and compliance frameworks. - Experience implementing monitoring, alerting, and observability strategies across distributed workloads. - Strong troubleshooting, communication, and documentation skills. Requirements - Design and implement enterprise-grade Azure cloud architectures spanning compute, networking, storage, identity, and data services, with explicit attention to scalability, security, and total cost of ownership. - Develop, maintain, and continuously improve infrastructure-as-code using Terraform, Bicep, or ARM templates. - Configure and manage Azure landing zones, virtual networks, subnets, route tables, and network security groups. - Implement secure identity, access management, and governance controls using Azure Active Directory. - Architect and operate Azure Kubernetes Service (AKS) clusters. - Deploy, scale, and tune Azure data and analytics platforms. - Build and operate comprehensive CI/CD pipelines using Azure DevOps or GitHub Actions. - Design and implement robust observability practices using Azure Monitor and other tools. - Drive Azure cost optimization initiatives. - Implement disaster-recovery and business-continuity strategies. - Strengthen security posture by integrating Microsoft Defender for Cloud and other security tools. - Collaborate closely with application teams to architect cloud-native solutions. - Develop automation scripts and tooling in PowerShell, Bash, and Python. - Mentor junior engineers and lead architecture reviews. Benefits - 100% Remote (U.S.) - Full-time, Direct W2 - Salary Range: $100,000–$150,000 Annually Company Description Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Network Engineer – Office Infrastructure
Freedom24Freedom24 is a platform for online trading of stocks, ETFs, bonds on major stock exchanges in the US, Europe and Asia
• Administer and develop the office network infrastructure (primarily MikroTik, with Ubiquiti deployed in several locations). • Reconfigure network segmentation across all office locations, optimize the existing infrastructure, and eliminate bottlenecks. • Design and implement high-availability solutions for network equipment and WAN connectivity. • Deploy network monitoring systems and introduce additional services to simplify network administration. • Maintain and update the corporate knowledge base (Wiki), keep network diagrams for major office locations current, and manage the organization's IP address space (IPAM). • Remotely prepare and configure network equipment for new office openings. • Work closely with Internet service providers (ISPs) and the team responsible for the data center network core. • Diagnose and resolve incidents related to network infrastructure and associated services.
Senior Platform Engineer, Network Infrastructure
NVIDIABased in Santa Clara, California, with additional offices throughout the U.S., South America, and Canada, NVIDIA is committed to fostering a work environment wh
• Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments. • Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery. • Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi-cluster delivery through GitOps. • Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features. • Diagnose complex Kubernetes platform and hosted-service failures involving control-plane health, cluster networking, storage, scheduling, workload placement, and multi-cluster dependencies. • Drive issues from initial signal through verified resolution. • Define production-readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery. • Participate in CFR’s production on-call rotation, including scheduled after-hours and weekend coverage. • Lead incident response and recovery, then drive corrective actions to completion.
Cloud Infrastructure Engineer
SupplyHouse.comPlumbing, Heating & HVAC Supplies. Real People. Real Service.
Role Description Through an Employer of Record (EOR), we are looking for a new Cloud Infrastructure Engineer in India to join our growing IT Team. This individual will report into our VP, IT and will support, maintain, and enhance our cloud and production infrastructure environments with a strong focus on Google Cloud Platform (GCP), Kubernetes operations, infrastructure automation, and platform reliability. If you are experienced in cloud infrastructure operations, Kubernetes administration, infrastructure automation, CI/CD support, and production incident management, we’d love to hear from you! Role Type: Full-Time Location: Remote from India Schedule: Monday through Friday with a minimum schedule overlap of 4-5 hours per day with 8:00 a.m. to 5:00 p.m. U.S. Eastern Time to ensure effective collaboration Base Salary: $20,258 - $25,660 USD per year Responsibilities - Cloud Infrastructure Management - Manage/maintain production infrastructure across dev, staging, and production environments - Support deployment, configuration, scaling, and lifecycle management of cloud services - Monitor system health, performance, availability, and capacity - Handle patching, upgrades, and maintenance activities - Ensure infrastructure reliability, uptime, and operational stability - Kubernetes & Containers - Support Kubernetes-based workloads and cluster operations - Assist with container orchestration, deployments, scaling, and troubleshooting - Maintain operational standards/best practices for containerized infrastructure - Partner with dev teams on application deployments and environment issues - Automation & CI/CD - Develop and maintain automation scripts/tooling for operational efficiency - Support infrastructure automation and provisioning via IaC - Assist with CI/CD pipeline integrations and deployment automation - Contribute to continuous improvement for scalability and reliability - Security & Access Management - Manage access control, IAM configs, and user permissions - Support security best practices and compliance requirements - Assist with vulnerability remediation and infrastructure hardening - Maintain auditability and operational governance standards - Incident Response & Support - Perform incident response, troubleshooting, root cause analysis - Participate in on-call rotations - Collaborate with internal teams on infrastructure/platform/application issues - Develop and maintain runbooks, docs, and support procedures - Database Operations - Manage backup, restore, recovery, and maintenance - Perform patching/upgrades/lifecycle management for MariaDB - Monitor availability, performance, and data integrity - Troubleshoot connectivity/performance issues with app teams Qualifications - High-level proficiency of written and verbal communication in English - 3–5 years of experience in cloud infrastructure, systems engineering, platform engineering, or infrastructure operations - Hands-on experience supporting Google Cloud Platform (GCP) environments - Experience with Kubernetes administration and containerized workloads - Strong understanding of Linux systems administration and infrastructure troubleshooting - Experience with Infrastructure as Code tools such as Terraform - Experience supporting CI/CD pipelines and deployment automation workflows - Familiarity with cloud networking concepts including VPCs, DNS, load balancing, and firewalls - Experience with monitoring, logging, and alerting platforms - Working knowledge of IAM, access control, and infrastructure security best practices - Experience supporting production systems and incident management processes - Experience with MariaDB or similar relational database platforms - Google Cloud certifications such as: Associate Cloud Engineer, Professional Cloud Architect, Professional Cloud DevOps Engineer - Experience with GitHub Actions, GitLab CI, Jenkins, or Cloud Build - Familiarity with observability tools such as Prometheus, Grafana, Datadog, or Splunk - Experience with scripting and automation using Python, Bash, or PowerShell - Exposure to SRE principles and operational excellence practices - Experience supporting hybrid cloud or multi-cloud environments Benefits - Comprehensive and affordable medical, dental, vision, and life insurance options - Competitive Provident Fund contributions - Paid time off and holidays - Mental health support and wellbeing program - Company-provided equipment and one-time $250 USD work from home stipend - $750 USD annual professional development budget - Company rewards and recognition program - And more!



