Strategic open source infrastructure for containers and virtual machines.
Senior AI Infrastructure, Platform Operations Engineer
Location
United States
Posted
2 days ago
Salary
0
Seniority
Senior
Job Description
Senior AI Infrastructure, Platform Operations Engineer
Mirantis
• Lead the investigation and resolution of complex infrastructure, networking, and platform-related incidents • Act as a senior escalation point for operational teams during critical service-impacting events • Support large-scale NVIDIA GPU infrastructure and high-performance networking environments • Troubleshoot complex Linux, Kubernetes, networking, storage, and hardware-related issues • Analyze platform performance, capacity, stability, and reliability trends to proactively identify risks • Lead root cause analysis activities and drive long-term corrective actions • Collaborate with engineering teams, hardware vendors, and datacenter personnel to resolve complex technical challenges • Participate in major incident management and service restoration activities • Provide technical leadership for Kubernetes platform operations and supporting infrastructure services • Drive improvements in platform reliability, observability, monitoring, and operational processes • Identify opportunities to automate repetitive operational activities and improve operational efficiency • Contribute to operational readiness reviews, infrastructure changes, upgrades, and service introductions • Support the adoption and operation of AI-powered infrastructure services and operational capabilities through k0rdent AI • Mentor and support AI Infrastructure & Platform Operations Engineers.
Job Requirements
- 7+ years of experience in infrastructure operations, platform operations, site reliability engineering, network operations, cloud operations, datacenter operations, or related technical roles
- Expert-level Linux administration and troubleshooting skills
- Strong networking expertise, including experience diagnosing complex performance, connectivity, and reliability issues
- Strong experience operating Kubernetes in production environments
- Experience supporting large-scale production infrastructure and distributed systems
- Proven experience leading technical investigations and managing complex incidents
- Experience performing root cause analysis and driving long-term operational improvements
- Strong understanding of observability, monitoring, and service reliability practices
- Excellent troubleshooting and analytical skills across multiple infrastructure domains
- Strong communication, collaboration, and stakeholder management skills.
Benefits
- Work with an established Silicon Valley leader in the cloud infrastructure industry
- Work with exceptionally passionate, talented and engaging colleagues
- Professional development and training
- Attend conferences and working groups
- Company outings, happy hours, hackathons, and tech talks
- Receive a competitive compensation package with a strong benefits plan
Related Guides
Related Categories
Related Job Pages
More Infrastructure Engineer Jobs
AI Infrastructure & Platform Operations Engineer
MirantisStrategic open source infrastructure for containers and virtual machines.
• Monitor, operate, and support production AI infrastructure platforms • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents • Support NVIDIA GPU infrastructure and associated platform services • Monitor and troubleshoot Kubernetes-based environments • Investigate performance, availability, and reliability issues across infrastructure and platform components • Collaborate with engineering teams, hardware vendors, Data Center personnel, and service delivery teams to resolve technical issues • Participate in incident response, root cause analysis, and operational improvement activities • Contribute to improvements in monitoring, observability, automation, and operational processes • Maintain operational documentation, runbooks, and knowledge articles
Azure Infrastructure Engineer
Bright Vision TechnologiesBright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.
Role Description We are seeking a skilled Azure Infrastructure Engineer to design, deploy, and operate large-scale, secure, and resilient cloud platforms on Microsoft Azure. In this role you will own the end-to-end cloud engineering lifecycle, including: - Architecture - Infrastructure-as-code - Automation - Security hardening - Cost optimization - Observability - Ongoing operational excellence for production workloads The ideal candidate will combine deep technical expertise across Azure services with strong DevOps engineering practices, and will partner closely with application development, security, and SRE teams to deliver cloud-native solutions that meet demanding business requirements for scalability, reliability, and compliance. Qualifications - Bachelor’s degree in Computer Science, Engineering, or a related technical discipline. - Five or more years of cloud engineering experience, with at least three years focused on Microsoft Azure in production environments. - Strong hands-on experience with Azure core services, including compute, storage, networking, identity, and platform-as-a-service offerings. - Production-level experience with infrastructure-as-code tools such as Terraform, Bicep, or ARM templates. - Solid experience designing and operating Azure Kubernetes Service (AKS) clusters at scale. - Hands-on experience with Azure DevOps or GitHub Actions for CI/CD across infrastructure and applications. - Strong scripting skills in PowerShell, Bash, and Python, with the ability to write maintainable automation code. - Deep understanding of cloud security principles, identity management, and compliance frameworks. - Experience implementing monitoring, alerting, and observability strategies across distributed workloads. - Strong troubleshooting, communication, and documentation skills. Requirements - Design and implement enterprise-grade Azure cloud architectures spanning compute, networking, storage, identity, and data services, with explicit attention to scalability, security, and total cost of ownership. - Develop, maintain, and continuously improve infrastructure-as-code using Terraform, Bicep, or ARM templates. - Configure and manage Azure landing zones, virtual networks, subnets, route tables, and network security groups. - Implement secure identity, access management, and governance controls using Azure Active Directory. - Architect and operate Azure Kubernetes Service (AKS) clusters. - Deploy, scale, and tune Azure data and analytics platforms. - Build and operate comprehensive CI/CD pipelines using Azure DevOps or GitHub Actions. - Design and implement robust observability practices using Azure Monitor and other tools. - Drive Azure cost optimization initiatives. - Implement disaster-recovery and business-continuity strategies. - Strengthen security posture by integrating Microsoft Defender for Cloud and other security tools. - Collaborate closely with application teams to architect cloud-native solutions. - Develop automation scripts and tooling in PowerShell, Bash, and Python. - Mentor junior engineers and lead architecture reviews. Benefits - 100% Remote (U.S.) - Full-time, Direct W2 - Salary Range: $100,000–$150,000 Annually Company Description Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Network Engineer – Office Infrastructure
Freedom24Freedom24 is a platform for online trading of stocks, ETFs, bonds on major stock exchanges in the US, Europe and Asia
• Administer and develop the office network infrastructure (primarily MikroTik, with Ubiquiti deployed in several locations). • Reconfigure network segmentation across all office locations, optimize the existing infrastructure, and eliminate bottlenecks. • Design and implement high-availability solutions for network equipment and WAN connectivity. • Deploy network monitoring systems and introduce additional services to simplify network administration. • Maintain and update the corporate knowledge base (Wiki), keep network diagrams for major office locations current, and manage the organization's IP address space (IPAM). • Remotely prepare and configure network equipment for new office openings. • Work closely with Internet service providers (ISPs) and the team responsible for the data center network core. • Diagnose and resolve incidents related to network infrastructure and associated services.
Senior Platform Engineer, Network Infrastructure
NVIDIANVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
• Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments. • Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery. • Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi-cluster delivery through GitOps. • Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features. • Diagnose complex Kubernetes platform and hosted-service failures involving control-plane health, cluster networking, storage, scheduling, workload placement, and multi-cluster dependencies. • Drive issues from initial signal through verified resolution. • Define production-readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery. • Participate in CFR’s production on-call rotation, including scheduled after-hours and weekend coverage. • Lead incident response and recovery, then drive corrective actions to completion.


