
Mirantis
Remote Jobs
Strategic open source infrastructure for containers and virtual machines.
102 Jobs
Senior HPC Networking Engineer
MirantisStrategic open source infrastructure for containers and virtual machines.
Role Description We are seeking a highly skilled Senior HPC Networking Engineer to design, deploy, manage, and troubleshoot high-performance networking environments. The ideal candidate will have deep expertise in InfiniBand technologies, strong general networking knowledge, and hands-on experience with Fortinet solutions. You will play a critical role in ensuring the performance, reliability, and scalability of HPC infrastructure. Key Responsibilities - Design, deploy, and maintain high-performance network infrastructures for HPC environments, with a strong focus on InfiniBand fabrics. - Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance. - Manage and optimize InfiniBand components, including switches, HCAs, subnet managers, and fabric configurations. - Perform performance tuning, monitoring, and capacity planning for HPC networking systems. - Implement and maintain network security using Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer). - Diagnose and resolve issues related to routing, switching, latency, and throughput across hybrid network environments. - Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations. - Develop and maintain documentation for network architecture, configurations, and operational procedures. - Participate in on-call rotations and provide escalation support for critical incidents. - Lead or contribute to network upgrades, migrations, and new deployments. Qualifications - 5+ years of experience in network engineering, with a focus on HPC or data center environments. - Strong hands-on experience with InfiniBand technologies (e.g., Mellanox/NVIDIA). - Solid understanding of networking fundamentals: TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network design. - Proven experience deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies). - Experience with network performance analysis and troubleshooting tools. - Familiarity with Linux systems and scripting for automation (e.g., Bash, Python). - Strong analytical and problem-solving skills. Preferred Qualifications - Experience with large-scale HPC clusters or AI/ML infrastructure. - Knowledge of RDMA, MPI, and low-latency networking concepts. - Certifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent. - Experience with automation and Infrastructure as Code tools (e.g., Ansible, Terraform). Soft Skills - Strong communication and collaboration skills. - Ability to work independently and handle complex technical challenges. - Detail-oriented with a proactive approach to problem-solving. Benefits - Opportunity to work on cutting-edge HPC infrastructure. - Collaborative and innovative work environment. - Competitive salary and benefits package. - Professional development and training. - Attend conferences and working groups. - Company outings, happy hours, hackathons, and tech talks. - Receive a competitive compensation package with a strong benefits plan.
Senior AI Infrastructure, Platform Operations Engineer
MirantisStrategic open source infrastructure for containers and virtual machines.
• Lead the investigation and resolution of complex infrastructure, networking, and platform-related incidents • Act as a senior escalation point for operational teams during critical service-impacting events • Support large-scale NVIDIA GPU infrastructure and high-performance networking environments • Troubleshoot complex Linux, Kubernetes, networking, storage, and hardware-related issues • Analyze platform performance, capacity, stability, and reliability trends to proactively identify risks • Lead root cause analysis activities and drive long-term corrective actions • Collaborate with engineering teams, hardware vendors, and datacenter personnel to resolve complex technical challenges • Participate in major incident management and service restoration activities • Provide technical leadership for Kubernetes platform operations and supporting infrastructure services • Drive improvements in platform reliability, observability, monitoring, and operational processes • Identify opportunities to automate repetitive operational activities and improve operational efficiency • Contribute to operational readiness reviews, infrastructure changes, upgrades, and service introductions • Support the adoption and operation of AI-powered infrastructure services and operational capabilities through k0rdent AI • Mentor and support AI Infrastructure & Platform Operations Engineers.
AI Infrastructure & Platform Operations Engineer
MirantisStrategic open source infrastructure for containers and virtual machines.
• Monitor, operate, and support production AI infrastructure platforms • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents • Support NVIDIA GPU infrastructure and associated platform services • Monitor and troubleshoot Kubernetes-based environments • Investigate performance, availability, and reliability issues across infrastructure and platform components • Collaborate with engineering teams, hardware vendors, Data Center personnel, and service delivery teams to resolve technical issues • Participate in incident response, root cause analysis, and operational improvement activities • Contribute to improvements in monitoring, observability, automation, and operational processes • Maintain operational documentation, runbooks, and knowledge articles
Senior Automation Software Tester
MirantisStrategic open source infrastructure for containers and virtual machines.
• Testing Strategy & Leadership: Contribute to driving the overall QA strategy for Mirantis Secure Registry, Kubernetes and cloud-native environments, aligning quality goals with business objectives and release deadlines. • Evaluate and champion best-in-class testing practices and frameworks for distributed systems. • Lead risk analysis and test planning for large-scale, multi-cloud registry, infrastructure and container orchestration. • Architect test frameworks and infrastructure for validating microservices and infrastructure components in multi-cluster and hybrid-cloud environments. • Oversee the design of complex test scenarios simulating production-like workloads, resource scaling, failure injection, and recovery across distributed clusters. • Spearhead the development of scalable and maintainable test automation integrated with CI/CD (Jenkins, GitHub Actions, etc.). • Leverage Kubernetes APIs, Helm, and service mesh tools to build comprehensive automation coverage. • Promote test infrastructure-as-code and drive IaC forward on the team making sure the infrastructure code is repeatable, extensible and reliable. • Mentor QA engineers and developers in advanced testing techniques such as risk-based testing, chaos engineering, performance and load testing. • Serve as the QA authority in design reviews, production readiness assessments, and incident retrospectives (escaped defects). • Collaborate with devops engineers to refine monitoring, alerting, and debugging strategies for testing automation being run in CI/CD. • Establish quality gates, release readiness metrics, and data driven feedback loops to ensure released software is production quality. • Drive initiatives to integrate performance, security, and chaos testing into the development lifecycle. • Advocate for a culture of quality across teams and influence architecture decisions with testing in mind.
Senior Automation Software Tester, MSR Team
MirantisStrategic open source infrastructure for containers and virtual machines.
• Testing Strategy & Leadership: Contribute to driving the overall QA strategy for Mirantis Secure Registry, Kubernetes and cloud-native environments, aligning quality goals with business objectives and release deadlines. • Evaluate and champion best-in-class testing practices and frameworks for distributed systems. • Lead risk analysis and test planning for large-scale, multi-cloud registry, infrastructure and container orchestration. • Architecting Test Systems: Architect test frameworks and infrastructure for validating microservices and infrastructure components in multi-cluster and hybrid-cloud environments. • Oversee the design of complex test scenarios simulating production-like workloads, resource scaling, failure injection, and recovery across distributed clusters. • Automation & Scalability: Spearhead the development of scalable and maintainable test automation integrated with CI/CD (Jenkins, GitHub Actions, etc.). • Leverage Kubernetes APIs, Helm, and service mesh tools to build comprehensive automation coverage, including system health, failover behavior, and network resilience. • Promote test infrastructure-as-code and drive IaC forward on the team making sure the infrastructure code is repeatable, extensible and reliable. • Mentorship & Technical Influence: Mentor QA engineers and developers in advanced testing techniques such as: risk-based testing, chaos engineering, performance and load testing etc. • Serve as the QA authority in design reviews, production readiness assessments, and incident retrospectives (escaped defects). • Collaborate with devops engineers to refine monitoring, alerting, and debugging strategies for testing automation being run in CI/CD. • Quality Advocacy & Continuous Improvement: Establish quality gates, release readiness metrics, and data driven feedback loops to ensure released software is production quality. • Drive initiatives to integrate performance, security, and chaos testing into the development lifecycle. • Advocate for a culture of quality across teams and influence architecture decisions with testing in mind.
Technical Product Manager, AI Cloud Networking
MirantisStrategic open source infrastructure for containers and virtual machines.
• Own the vision, roadmap, and priorities for k0rdent AI networking, defining how customers design, automate, and operate every aspect of their network, including underlay fabric management, tenant connectivity, RDMA, DNS/IPAM, and more • Translate requirements from NeoClouds, GPU clouds, telcos, sovereign clouds, and enterprise platform teams into clear product direction • Partner with engineering and architecture to define requirements, evaluate trade-offs, and ship secure, scalable, reliable networking capabilities • Manage the networking backlog, using feedback from production deployments and design partners to refine roadmap priorities and positioning • Define positioning, packaging, and competitive differentiation for k0rdent AI networking • Create field-facing assets, including technical briefs, battlecards, and reference architectures; support strategic accounts as the networking product lead • Represent Mirantis at events, analyst briefings, and customer advisory boards; engage silicon, fabric, and ecosystem partners on reference architecture alignment
Principal HPC Network Engineer
MirantisStrategic open source infrastructure for containers and virtual machines.
• Design, deploy, and maintain high-performance network infrastructures for HPC environments, with a strong focus on InfiniBand fabrics. • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance. • Manage and optimize InfiniBand components, including switches, HCAs, subnet managers, and fabric configurations. • Perform performance tuning, monitoring, and capacity planning for HPC networking systems. • Implement and maintain network security using Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer). • Diagnose and resolve issues related to routing, switching, latency, and throughput across hybrid network environments. • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations. • Develop and maintain documentation for network architecture, configurations, and operational procedures. • Participate in on-call rotations and provide escalation support for critical incidents. • Lead or contribute to network upgrades, migrations, and new deployments.
Principal HPC Network Engineer
MirantisStrategic open source infrastructure for containers and virtual machines.
• Design, deploy, and maintain high-performance network infrastructures for HPC environments, with a strong focus on InfiniBand fabrics. • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance. • Manage and optimize InfiniBand components, including switches, HCAs, subnet managers, and fabric configurations. • Perform performance tuning, monitoring, and capacity planning for HPC networking systems. • Implement and maintain network security using Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer). • Diagnose and resolve issues related to routing, switching, latency, and throughput across hybrid network environments. • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations. • Develop and maintain documentation for network architecture, configurations, and operational procedures. • Participate in on-call rotations and provide escalation support for critical incidents. • Lead or contribute to network upgrades, migrations, and new deployments.
Principal HPC Network Engineer, Remote in the EU
MirantisStrategic open source infrastructure for containers and virtual machines.
• Design, deploy, and maintain high-performance network infrastructures for HPC environments, with a strong focus on InfiniBand fabrics. • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance. • Manage and optimize InfiniBand components, including switches, HCAs, subnet managers, and fabric configurations. • Perform performance tuning, monitoring, and capacity planning for HPC networking systems. • Implement and maintain network security using Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer). • Diagnose and resolve issues related to routing, switching, latency, and throughput across hybrid network environments. • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations. • Develop and maintain documentation for network architecture, configurations, and operational procedures. • Participate in on-call rotations and provide escalation support for critical incidents. • Lead or contribute to network upgrades, migrations, and new deployments.
Senior AI Infrastructure, Platform Operations Engineer
MirantisStrategic open source infrastructure for containers and virtual machines.
• Lead the investigation and resolution of complex infrastructure, networking, and platform-related incidents. • Act as a senior escalation point for operational teams during critical service-impacting events. • Support large-scale NVIDIA GPU infrastructure and high-performance networking environments. • Troubleshoot complex Linux, Kubernetes, networking, storage, and hardware-related issues. • Analyze platform performance, capacity, stability, and reliability trends to proactively identify risks. • Lead root cause analysis activities and drive long-term corrective actions. • Collaborate with engineering teams, hardware vendors, and datacenter personnel to resolve complex technical challenges. • Provide technical leadership for Kubernetes platform operations and supporting infrastructure services. • Drive improvements in platform reliability, observability, monitoring, and operational processes. • Identify opportunities to automate repetitive operational activities and improve operational efficiency. • Mentor and support AI Infrastructure & Platform Operations Engineers.
92more opportunities are still waiting for you.Log in now and take your next shot before someone else does.