This opportunity is available through a leading AI-driven work platform.
ML Infrastructure & Kernel Optimization Engineer
Location
United States
Posted
7 days ago
Salary
$65 - $105 / hour
Seniority
Mid Level
Job Description
ML Infrastructure & Kernel Optimization Engineer
24-MAG
Role Description We are sharing a specialised full-time consulting opportunity for US-based MLOps and ML systems engineers with production experience in JAX, PyTorch, distributed training infrastructure, and custom GPU kernel development using Pallas or Triton. This role supports a high-impact generative AI initiative focused on developing and evaluating advanced ML infrastructure tasks for frontier model training. Selected engineers will: - Design technically challenging problems - Produce rigorous solutions - Assess model-generated outputs - Help establish evaluation standards across training pipelines, distributed systems, framework-level optimisation, and GPU kernel performance Qualifications - At least 2 years of dedicated professional experience in MLOps, ML infrastructure, or ML systems engineering - Production experience with JAX, PyTorch, or both at meaningful scale - Hands-on experience writing or optimising custom GPU kernels using Pallas or Triton - Strong knowledge of model-training pipelines, distributed systems, accelerators, and performance optimisation - Experience working within a recognised technology, AI research, or high-performance engineering organisation - Demonstrable professional growth and increasing technical responsibility - Strong written communication and the ability to explain complex engineering decisions clearly - Reliable availability for a full-time, 40-hour weekday schedule Requirements - A degree in computer science, machine learning, electrical engineering, applied mathematics, or a related technical field is highly relevant - Graduate-level education in machine learning systems, distributed computing, or high-performance computing may be helpful - Equivalent professional experience in production ML infrastructure may also be considered - Advanced technical work involving GPU programming, compiler systems, or large-scale model training is especially valuable Benefits - Contribute to advanced generative AI and large-scale model-training initiatives - Apply deep expertise in JAX, PyTorch, Pallas, Triton, and ML infrastructure - Work on challenging problems spanning training systems, distributed computing, and GPU optimisation - Influence the quality of technical training data used in frontier AI development - Join a full-time remote engagement with competitive hourly compensation Contract Details - Full-time W-2 contingent employment arrangement - Fully remote role available to candidates based in the United States - Expected commitment of 40 hours per week during weekdays - This engagement requires full professional availability without conflicting employment or external commitments - Competitive rates between $65–$105 per hour depending on expertise and project scope - Immediate availability is preferred - Work may include onboarding, technical calibration, and ongoing quality-review activities - Project scope and duration may be adjusted according to programme requirements and performance About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy
Related Guides
Related Categories
Related Job Pages
More Infrastructure Engineer Jobs
• AI Platform Engineering: Design, deploy, and operate enterprise AI/ML platforms. • Build self-service platforms for Data Scientists and ML Engineers. • Deploy and operate Kubeflow, MLflow, KServe, Ray, or similar AI platforms. • Design infrastructure supporting model training, experimentation, feature engineering, and inference. • Build highly available and scalable model serving infrastructure. • GPU Infrastructure: Design and operate GPU clusters for large-scale AI workloads. • Optimize GPU scheduling, utilization, sharing, autoscaling, and resource allocation. • Deploy and manage NVIDIA GPU Operator and GPU-enabled Kubernetes environments. • Optimize distributed GPU training performance across multi-node clusters. • Troubleshoot AI infrastructure performance bottlenecks. • MLOps & Platform Automation: Build CI/CD pipelines for ML workloads. • Automate AI infrastructure provisioning using Infrastructure as Code. • Implement monitoring and observability for GPU utilization, model serving, training jobs, and inference latency. • Collaborate closely with Data Science teams to improve platform usability, performance, and reliability.
Infrastructure Engineer
Voltage ParkVoltage Park is an equal opportunity employer and makes employment decisions on the basis of merit. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other characteristic under federal, state, or local law. If you require an accommodation during the job application process, please notify your recruiter.
Role Description Voltage Park is seeking an Infrastructure Engineer with a focus on Observability to join our Infrastructure Engineering team. Our engineers design and operate the systems that manage thousands of bare-metal servers, GPUs, and high-performance networks across multiple data centers. This role combines the breadth of a core infrastructure engineer with a specialty in observability and telemetry. You’ll design and operate metrics, logs, traces, and alerting pipelines that provide actionable insights for both internal teams and external customers — helping to ensure reliability and transparency at scale. This is a fully remote position, although candidates must be based in the continental United States. Unfortunately, we are unable to provide sponsorship for this role. Responsibilities - Design, build, and maintain observability platforms spanning metrics, logs, traces, and events. - Create dashboards and alerting for internal stakeholders (InfraOps, Engineering, Customer Success) and scoped visibility for external customers. - Ingest and correlate telemetry from GPUs, CPUs, networking (Ethernet & InfiniBand), containers, APIs, and BMC/Redfish. - Implement noise-resistant alerting pipelines that improve detection and reduce operational load. - Collaborate with infrastructure, platform, and customer-facing teams to embed observability into workflows. - Contribute to broader infrastructure engineering projects beyond observability. Qualifications - 8+ years in infrastructure engineering, SRE, or observability roles. - Strong experience with monitoring systems (Prometheus, Grafana, ELK, VictoriaMetrics, or similar). - Proficiency in Python, Go, or bash for automation and data integration. - Familiarity with container/Kubernetes observability. - Understanding of streaming telemetry pipelines (Kafka, OTEL, Promtail, or equivalent). - Strong written and verbal communication skills. Ideal Experiences - Experience with GPU observability, particularly NVIDIA DCGM. - Designing multi-tenant observability solutions with RBAC and scoped queries. - Prior work with correlation engines for RCA, forecasting, or predictive alerting. - Broader exposure to infrastructure domains (networking, storage, provisioning). Culture - You enjoy working with a small, highly motivated team. - You’re comfortable balancing autonomy with company-wide priorities. - You value clarity, documentation, and actionable insights in observability systems. - You’re excited to specialize in observability while contributing as a core infrastructure engineer. Company Description Voltage Park is an equal opportunity employer and makes employment decisions on the basis of merit. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other characteristic under federal, state, or local law. If you require an accommodation during the job application process, please notify your recruiter.
• Diseñar, desarrollar y administrar soluciones en Microsoft Azure. • Implementar arquitecturas escalables, seguras y de alto rendimiento. • Colaborar en equipos multiculturales para el desarrollo de soluciones innovadoras. • Participar activamente en la mejora de procesos y buenas prácticas.
Infrastructure Storage Engineer
Bright Vision TechnologiesBright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.
Role Description We are seeking an Infrastructure Storage Engineer with deep expertise across enterprise storage platforms — NetApp ONTAP, Pure Storage, and Ceph — to design, deploy, and operate the storage foundation that supports our compute, virtualization, database, and Kubernetes workloads. The role spans block, file, and object storage across data center, edge, and cloud environments, with strong attention to performance, data protection, and cost. The ideal candidate has operated heterogeneous storage estates at scale, understands the storage characteristics of diverse workloads, and brings strong automation discipline to storage engineering work. Key Responsibilities - Design and operate enterprise storage platforms across NetApp ONTAP, Pure Storage FlashArray/FlashBlade, and Ceph. - Implement and manage SAN, NAS, and object storage across data center and cloud environments. - Build storage solutions for VMware, Kubernetes (CSI drivers), and database workloads. - Design and operate replication, snapshot, and backup strategies for critical data assets. - Implement and operate cloud-tiered storage with FabricPool, CloudVolumes, or equivalent. - Build storage automation using Ansible, Terraform, REST APIs, and platform-native automation. - Operate Ceph clusters including OSD lifecycle, CRUSH maps, RGW, RBD, and CephFS. - Design and operate storage performance management for high-throughput and latency-sensitive workloads. - Implement data protection strategies including ransomware protection and immutable snapshots. - Build storage observability across capacity, performance, and health dimensions. - Lead storage upgrades, firmware lifecycle management, and security patching across the storage estate. - Drive storage cost optimization including deduplication, compression, and tiering strategies. - Troubleshoot complex storage issues spanning array, network, and host layers. - Stay current with storage industry developments and emerging platforms. Qualifications - Bachelor’s degree in Computer Science, Information Systems, or a related field. - Five or more years of enterprise storage engineering experience. - Deep expertise in NetApp ONTAP or Pure Storage FlashArray/FlashBlade. - Hands-on experience with Ceph at production scale. - Strong understanding of SAN, NAS, and object storage protocols. - Experience with storage integration for VMware and Kubernetes (CSI). - Strong scripting and automation skills using Ansible, Terraform, or REST APIs. - Strong understanding of storage performance, capacity planning, and cost optimization. - Strong troubleshooting skills across storage, network, and host layers. - Excellent communication and collaboration skills. Preferred Qualifications - Vendor certifications (NCDA, NCIE, Pure Storage certifications). - Experience with hybrid-cloud storage replication. - Familiarity with software-defined storage platforms beyond Ceph. - Experience with object storage at petabyte scale. - Exposure to AI/ML workloads with extreme I/O requirements. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3544. Learn more about Bright Vision Technologies at www.bvteck.com .

