Mozn logo
Mozn

AI Powerhouse of The Region

AI Infrastructure Engineer III

Infrastructure EngineerInfrastructure EngineerFull TimeRemoteSeniorTeam 201-500Since 2017H1B No SponsorCompany SiteLinkedIn

Location

Egypt

Posted

2 days ago

Salary

0

Seniority

Senior

Job Description

AI Infrastructure Engineer III

Mozn

• AI Platform Engineering: Design, deploy, and operate enterprise AI/ML platforms. • Build self-service platforms for Data Scientists and ML Engineers. • Deploy and operate Kubeflow, MLflow, KServe, Ray, or similar AI platforms. • Design infrastructure supporting model training, experimentation, feature engineering, and inference. • Build highly available and scalable model serving infrastructure. • GPU Infrastructure: Design and operate GPU clusters for large-scale AI workloads. • Optimize GPU scheduling, utilization, sharing, autoscaling, and resource allocation. • Deploy and manage NVIDIA GPU Operator and GPU-enabled Kubernetes environments. • Optimize distributed GPU training performance across multi-node clusters. • Troubleshoot AI infrastructure performance bottlenecks. • MLOps & Platform Automation: Build CI/CD pipelines for ML workloads. • Automate AI infrastructure provisioning using Infrastructure as Code. • Implement monitoring and observability for GPU utilization, model serving, training jobs, and inference latency. • Collaborate closely with Data Science teams to improve platform usability, performance, and reliability.

Job Requirements

  • 4-6 years of experience in AI Infrastructure, MLOps, Platform Engineering, or Cloud Engineering.
  • Strong hands-on experience with Kubernetes.
  • Experience with Kubeflow, MLflow, or similar ML platform technologies.
  • Experience operating GPU infrastructure for AI workloads.
  • Strong understanding of NVIDIA GPU technologies, CUDA fundamentals, and GPU optimization.
  • Experience supporting distributed training workloads.
  • Experience with model serving platforms such as KServe, Triton Inference Server, Ray Serve, or similar.
  • Experience with AWS, GCP, OCI, or Azure AI platforms.
  • Experience automating infrastructure using Terraform, Helm, GitOps, or Ansible.
  • Strong scripting or programming skills in Python, Bash, or Go.
  • Experience with Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, or equivalent observability platforms.
  • Preferred Qualifications: Experience with PyTorch, TensorFlow, Hugging Face, or JAX.
  • Experience with distributed training frameworks such as Ray, DeepSpeed, Horovod, or NCCL.
  • Experience with Vector Databases, LLM infrastructure, RAG architectures, or GenAI platforms.
  • Experience operating inference platforms for large language models.
  • Experience supporting AI research or Data Science teams in production environments.
  • Contributions to Cloud Native, Kubernetes, AI, or ML open-source communities.
  • Cloud, Kubernetes, NVIDIA, or AI/ML certifications are a plus.

Benefits

  • You will be at the forefront of an exciting time for the Middle East, joining a high-growth rocket-ship in an exciting space
  • You will be given a lot of responsibility and trust.
  • We believe that the best results come when the people responsible for a function are given the freedom to do what they think is best
  • The fundamentals will be taken care of: competitive compensation, top-tier health insurance, and an enabling culture so that you can focus on what you do best
  • You will enjoy a fun and dynamic workplace working alongside some of the greatest minds in AI
  • We believe strength lies in difference, embracing all for who they are and empowered to be the best version of themselves

Related Categories

Related Job Pages

More Infrastructure Engineer Jobs

Infrastructure Engineer

Voltage Park

Voltage Park is an equal opportunity employer and makes employment decisions on the basis of merit. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other characteristic under federal, state, or local law. If you require an accommodation during the job application process, please notify your recruiter.

Role Description Voltage Park is seeking an Infrastructure Engineer with a focus on Observability to join our Infrastructure Engineering team. Our engineers design and operate the systems that manage thousands of bare-metal servers, GPUs, and high-performance networks across multiple data centers. This role combines the breadth of a core infrastructure engineer with a specialty in observability and telemetry. You’ll design and operate metrics, logs, traces, and alerting pipelines that provide actionable insights for both internal teams and external customers — helping to ensure reliability and transparency at scale. This is a fully remote position, although candidates must be based in the continental United States. Unfortunately, we are unable to provide sponsorship for this role. Responsibilities - Design, build, and maintain observability platforms spanning metrics, logs, traces, and events. - Create dashboards and alerting for internal stakeholders (InfraOps, Engineering, Customer Success) and scoped visibility for external customers. - Ingest and correlate telemetry from GPUs, CPUs, networking (Ethernet & InfiniBand), containers, APIs, and BMC/Redfish. - Implement noise-resistant alerting pipelines that improve detection and reduce operational load. - Collaborate with infrastructure, platform, and customer-facing teams to embed observability into workflows. - Contribute to broader infrastructure engineering projects beyond observability. Qualifications - 8+ years in infrastructure engineering, SRE, or observability roles. - Strong experience with monitoring systems (Prometheus, Grafana, ELK, VictoriaMetrics, or similar). - Proficiency in Python, Go, or bash for automation and data integration. - Familiarity with container/Kubernetes observability. - Understanding of streaming telemetry pipelines (Kafka, OTEL, Promtail, or equivalent). - Strong written and verbal communication skills. Ideal Experiences - Experience with GPU observability, particularly NVIDIA DCGM. - Designing multi-tenant observability solutions with RBAC and scoped queries. - Prior work with correlation engines for RCA, forecasting, or predictive alerting. - Broader exposure to infrastructure domains (networking, storage, provisioning). Culture - You enjoy working with a small, highly motivated team. - You’re comfortable balancing autonomy with company-wide priorities. - You value clarity, documentation, and actionable insights in observability systems. - You’re excited to specialize in observability while contributing as a core infrastructure engineer. Company Description Voltage Park is an equal opportunity employer and makes employment decisions on the basis of merit. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other characteristic under federal, state, or local law. If you require an accommodation during the job application process, please notify your recruiter.

United States
Full TimeRemoteTeam 1,001-5,000Since 1977H1B No Sponsor

• Diseñar, desarrollar y administrar soluciones en Microsoft Azure. • Implementar arquitecturas escalables, seguras y de alto rendimiento. • Colaborar en equipos multiculturales para el desarrollo de soluciones innovadoras. • Participar activamente en la mejora de procesos y buenas prácticas.

Colombia

Infrastructure Storage Engineer

Bright Vision Technologies

Bright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.

Role Description We are seeking an Infrastructure Storage Engineer with deep expertise across enterprise storage platforms — NetApp ONTAP, Pure Storage, and Ceph — to design, deploy, and operate the storage foundation that supports our compute, virtualization, database, and Kubernetes workloads. The role spans block, file, and object storage across data center, edge, and cloud environments, with strong attention to performance, data protection, and cost. The ideal candidate has operated heterogeneous storage estates at scale, understands the storage characteristics of diverse workloads, and brings strong automation discipline to storage engineering work. Key Responsibilities - Design and operate enterprise storage platforms across NetApp ONTAP, Pure Storage FlashArray/FlashBlade, and Ceph. - Implement and manage SAN, NAS, and object storage across data center and cloud environments. - Build storage solutions for VMware, Kubernetes (CSI drivers), and database workloads. - Design and operate replication, snapshot, and backup strategies for critical data assets. - Implement and operate cloud-tiered storage with FabricPool, CloudVolumes, or equivalent. - Build storage automation using Ansible, Terraform, REST APIs, and platform-native automation. - Operate Ceph clusters including OSD lifecycle, CRUSH maps, RGW, RBD, and CephFS. - Design and operate storage performance management for high-throughput and latency-sensitive workloads. - Implement data protection strategies including ransomware protection and immutable snapshots. - Build storage observability across capacity, performance, and health dimensions. - Lead storage upgrades, firmware lifecycle management, and security patching across the storage estate. - Drive storage cost optimization including deduplication, compression, and tiering strategies. - Troubleshoot complex storage issues spanning array, network, and host layers. - Stay current with storage industry developments and emerging platforms. Qualifications - Bachelor’s degree in Computer Science, Information Systems, or a related field. - Five or more years of enterprise storage engineering experience. - Deep expertise in NetApp ONTAP or Pure Storage FlashArray/FlashBlade. - Hands-on experience with Ceph at production scale. - Strong understanding of SAN, NAS, and object storage protocols. - Experience with storage integration for VMware and Kubernetes (CSI). - Strong scripting and automation skills using Ansible, Terraform, or REST APIs. - Strong understanding of storage performance, capacity planning, and cost optimization. - Strong troubleshooting skills across storage, network, and host layers. - Excellent communication and collaboration skills. Preferred Qualifications - Vendor certifications (NCDA, NCIE, Pure Storage certifications). - Experience with hybrid-cloud storage replication. - Familiarity with software-defined storage platforms beyond Ceph. - Experience with object storage at petabyte scale. - Exposure to AI/ML workloads with extreme I/O requirements. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3544. Learn more about Bright Vision Technologies at www.bvteck.com .

United States
$100K - $150K / year
Mirantis logo

Senior AI Infrastructure, Platform Operations Engineer

Mirantis

Strategic open source infrastructure for containers and virtual machines.

Full TimeRemoteTeam 501-1,000H1B Sponsor

• Lead the investigation and resolution of complex infrastructure, networking, and platform-related incidents • Act as a senior escalation point for operational teams during critical service-impacting events • Support large-scale NVIDIA GPU infrastructure and high-performance networking environments • Troubleshoot complex Linux, Kubernetes, networking, storage, and hardware-related issues • Analyze platform performance, capacity, stability, and reliability trends to proactively identify risks • Lead root cause analysis activities and drive long-term corrective actions • Collaborate with engineering teams, hardware vendors, and datacenter personnel to resolve complex technical challenges • Participate in major incident management and service restoration activities • Provide technical leadership for Kubernetes platform operations and supporting infrastructure services • Drive improvements in platform reliability, observability, monitoring, and operational processes • Identify opportunities to automate repetitive operational activities and improve operational efficiency • Contribute to operational readiness reviews, infrastructure changes, upgrades, and service introductions • Support the adoption and operation of AI-powered infrastructure services and operational capabilities through k0rdent AI • Mentor and support AI Infrastructure & Platform Operations Engineers.

United States