TensorWave logo
TensorWave

GPU poor? Contact us for your AI cloud compute needs!

Infrastructure Engineer – Storage Platform

Infrastructure EngineerInfrastructure EngineerFull TimeRemoteSeniorTeam 11-50H1B No SponsorCompany SiteLinkedIn

Location

United States

Posted

5 days ago

Salary

0

Seniority

Senior

Bachelor Degree4 yrs expEnglishAnsibleKubernetesLinuxTerraform

Job Description

Infrastructure Engineer – Storage Platform

TensorWave

• Operate and maintain distributed storage platforms, including Ceph (RBD, CephFS, RGW), High-performance NAS platforms (e.g., Weka, VAST Data) • Manage storage lifecycle operations - cluster expansion, upgrades and migrations • Monitor and maintain storage health, including capacity utilization, data distribution and balance, cluster state and recovery operations • Analyze and troubleshoot storage performance across IOPS, throughput, and latency (including tail latency) • Identify and remediate bottlenecks across disk subsystems, network paths (including RDMA where applicable), client access patterns • Support incident response and root cause analysis for storage-related issues • Ensure storage platforms meet performance expectations for GPU and Kubernetes workloads • Operate and support Kubernetes-integrated storage - CSI drivers, StorageClasses, PersistentVolumes / PersistentVolumeClaims • Troubleshoot storage-related issues in Kubernetes environments, including stateful workloads, performance inconsistencies, scheduling and provisioning failures • Execute and improve automation for storage deployment and operations using Ansible, Terraform, Kubernetes manifests / Helm • Contribute to improving monitoring and alerting, operational workflows, runbooks and documentation • Partner with DevOps and Platform Engineering (automation and orchestration), Network Engineering (high-throughput and RDMA networking), Compute / Virtualization teams • Help ensure end-to-end performance across compute, network, and storage layers

Job Requirements

  • 4–7+ years of experience in infrastructure, systems, or storage operations
  • Strong hands-on experience operating distributed storage systems in production
  • Experience with Ceph (RBD, CephFS, or RGW)
  • Strong Linux systems knowledge
  • Experience with modern storage platforms such as:
  • Weka, VAST Data, or similar high-performance systems
  • Solid understanding of:
  • Storage performance characteristics (IOPS, throughput, latency)
  • Data replication and failure domains
  • Ability to troubleshoot across:
  • Storage systems
  • Network paths
  • Compute clients

Benefits

  • Stock Options
  • 100% paid Medical, Dental, and Vision insurance for Employees
  • Company Health Savings Account Contributions
  • 100% paid Short Term and Long Term Disability Insurance for Employees
  • Life and Voluntary Supplemental Insurance Options
  • Other Insurance Options, such as Pet & Legal Insurance
  • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support
  • Flexible Spending Account
  • 401(k)
  • Employee Assistance Program
  • Flexible PTO
  • Paid Holidays
  • Parental Leave
  • Other In-Office Perks

Related Categories

Related Job Pages

More Infrastructure Engineer Jobs

Full TimeRemoteTeam 501-1,000

• Manage and support infrastructure across Windows, Unix/Linux, networking, and cloud platforms • Monitor system alerts, triage issues, and respond quickly to maintain uptime • Lead and execute infrastructure projects, partnering with IT and business stakeholders • Support cross-functional teams with infrastructure-related incidents and escalations • Execute core processes including patching, monitoring, and system maintenance • Assist with M&A integrations, including system migrations and infrastructure support • Develop and maintain documentation (architecture, SOPs, standards, policies) • Continuously look for ways to improve performance, reliability, and scalability • Support additional initiatives and projects as needed

United States
Ravelin Technology logo

Infrastructure Engineer, Mid Level

Ravelin Technology

Make smarter decisions on fraud and payments.

Full TimeRemoteTeam 51-200Since 2014

- We are looking for someone to join our Infrastructure Engineering Team to work on the core infrastructure that underpins our platform. - You will work closely with teams across engineering and in the wider organisation to build systems with reliability, security, and scalability at their heart. - You care deeply about developer experience and will contribute to improving CI/CD pipelines, enabling new capabilities in GCP, and optimising performance of our Kubernetes cluster. - If you love automating everything, treating infrastructure as code, and building resilient systems, we want to hear from you.

United States
fal logo

Software Engineer, Infrastructure

fal

Generative media platform for developers.

Full TimeRemoteTeam 51-200

• Build and maintain Python fleet tracking system that manages the full lifecycle of servers including contracting and procurement, target use, pricing, availability, health, RMAs, etc • Build server management tooling that automates provisioning, health checks, GPU diagnostics, recovery and alerting • Create and maintain metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk failures, network issues, thermals) • Leverage AI to an extreme level to build tools and automate alerting and recovery • Implement and enforce OS-level security: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation • Manage and optimize distributed and local storage systems supporting model weights, checkpoints, and ephemeral scratch: NVMe arrays, NFS, parallel file systems, and object storage • Tune Linux systems for AI workloads: kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stack optimization (NVIDIA drivers, CUDA, container runtimes) • Develop a suite of automated error detection and recovery processes • Work with partners to solve technical issues

Turkey
Full TimeRemoteTeam 1,001-5,000Since 2016

• Define and maintain target-state architecture for IT and OT infrastructure, including networks, servers, virtualization, cloud platforms, identity, and supporting services. • Design secure, resilient, and scalable architectures for enterprise and plant environments, supporting availability, performance, and operational safety requirements. • Develop and maintain architecture standards, patterns, and reference designs for IT and OT environments. • Partner with enterprise architecture, application, and OT teams to ensure consistent adoption of standards and integration across domains. • Collaborate with OT teams to design and secure infrastructure supporting ICS, DCS, PLCs, SCADA, and plant network environments. • Support secure IT/OT convergence while respecting operational constraints, safety, and uptime requirements. • Ensure OT infrastructure designs align with industry best practices and regulatory expectations for industrial environments. • Embed security-by-design principles into all IT and OT infrastructure architectures. • Define and enforce cybersecurity architecture controls aligned to frameworks such as NIST CSF, NIST 800-53/82, IEC 62443, and Zero Trust principles. • Partner with cybersecurity teams to design controls for network segmentation, identity and access management, monitoring, logging, vulnerability management, and incident response. • Participate in architecture and security risk assessments; identify gaps and drive remediation through design and engineering standards. • Ensure infrastructure solutions comply with corporate policies, cybersecurity standards, and regulatory requirements. • Support audits, risk assessments, and tabletop exercises by providing architectural insight and documentation. • Maintain accurate architecture documentation, diagrams, and decision records. • Serve as a trusted technical advisor to IT, OT, cybersecurity, and business stakeholders. • Provide architectural guidance during project intake, design reviews, and technology evaluations. • Mentor infrastructure engineers and encourage knowledge transfer to reduce key-person dependencies. • Evaluate emerging technologies and recommend improvements to enhance security, resilience, and cost efficiency.

New Jersey + 2 moreAll locations: New Jersey | Pennsylvania | Virginia
$130.7K - $196.1K / year