GPU poor? Contact us for your AI cloud compute needs!
Infrastructure Engineer – Storage Platform
Location
United States
Posted
5 days ago
Salary
0
Seniority
Senior
Job Description
Infrastructure Engineer – Storage Platform
TensorWave
• Operate and maintain distributed storage platforms, including Ceph (RBD, CephFS, RGW), High-performance NAS platforms (e.g., Weka, VAST Data) • Manage storage lifecycle operations - cluster expansion, upgrades and migrations • Monitor and maintain storage health, including capacity utilization, data distribution and balance, cluster state and recovery operations • Analyze and troubleshoot storage performance across IOPS, throughput, and latency (including tail latency) • Identify and remediate bottlenecks across disk subsystems, network paths (including RDMA where applicable), client access patterns • Support incident response and root cause analysis for storage-related issues • Ensure storage platforms meet performance expectations for GPU and Kubernetes workloads • Operate and support Kubernetes-integrated storage - CSI drivers, StorageClasses, PersistentVolumes / PersistentVolumeClaims • Troubleshoot storage-related issues in Kubernetes environments, including stateful workloads, performance inconsistencies, scheduling and provisioning failures • Execute and improve automation for storage deployment and operations using Ansible, Terraform, Kubernetes manifests / Helm • Contribute to improving monitoring and alerting, operational workflows, runbooks and documentation • Partner with DevOps and Platform Engineering (automation and orchestration), Network Engineering (high-throughput and RDMA networking), Compute / Virtualization teams • Help ensure end-to-end performance across compute, network, and storage layers
Job Requirements
- 4–7+ years of experience in infrastructure, systems, or storage operations
- Strong hands-on experience operating distributed storage systems in production
- Experience with Ceph (RBD, CephFS, or RGW)
- Strong Linux systems knowledge
- Experience with modern storage platforms such as:
- Weka, VAST Data, or similar high-performance systems
- Solid understanding of:
- Storage performance characteristics (IOPS, throughput, latency)
- Data replication and failure domains
- Ability to troubleshoot across:
- Storage systems
- Network paths
- Compute clients
Benefits
- Stock Options
- 100% paid Medical, Dental, and Vision insurance for Employees
- Company Health Savings Account Contributions
- 100% paid Short Term and Long Term Disability Insurance for Employees
- Life and Voluntary Supplemental Insurance Options
- Other Insurance Options, such as Pet & Legal Insurance
- Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support
- Flexible Spending Account
- 401(k)
- Employee Assistance Program
- Flexible PTO
- Paid Holidays
- Parental Leave
- Other In-Office Perks
Related Guides
Related Categories
Related Job Pages
More Infrastructure Engineer Jobs
• Manage and support infrastructure across Windows, Unix/Linux, networking, and cloud platforms • Monitor system alerts, triage issues, and respond quickly to maintain uptime • Lead and execute infrastructure projects, partnering with IT and business stakeholders • Support cross-functional teams with infrastructure-related incidents and escalations • Execute core processes including patching, monitoring, and system maintenance • Assist with M&A integrations, including system migrations and infrastructure support • Develop and maintain documentation (architecture, SOPs, standards, policies) • Continuously look for ways to improve performance, reliability, and scalability • Support additional initiatives and projects as needed
- We are looking for someone to join our Infrastructure Engineering Team to work on the core infrastructure that underpins our platform. - You will work closely with teams across engineering and in the wider organisation to build systems with reliability, security, and scalability at their heart. - You care deeply about developer experience and will contribute to improving CI/CD pipelines, enabling new capabilities in GCP, and optimising performance of our Kubernetes cluster. - If you love automating everything, treating infrastructure as code, and building resilient systems, we want to hear from you.
• Build and maintain Python fleet tracking system that manages the full lifecycle of servers including contracting and procurement, target use, pricing, availability, health, RMAs, etc • Build server management tooling that automates provisioning, health checks, GPU diagnostics, recovery and alerting • Create and maintain metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk failures, network issues, thermals) • Leverage AI to an extreme level to build tools and automate alerting and recovery • Implement and enforce OS-level security: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation • Manage and optimize distributed and local storage systems supporting model weights, checkpoints, and ephemeral scratch: NVMe arrays, NFS, parallel file systems, and object storage • Tune Linux systems for AI workloads: kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stack optimization (NVIDIA drivers, CUDA, container runtimes) • Develop a suite of automated error detection and recovery processes • Work with partners to solve technical issues
• Define and maintain target-state architecture for IT and OT infrastructure, including networks, servers, virtualization, cloud platforms, identity, and supporting services. • Design secure, resilient, and scalable architectures for enterprise and plant environments, supporting availability, performance, and operational safety requirements. • Develop and maintain architecture standards, patterns, and reference designs for IT and OT environments. • Partner with enterprise architecture, application, and OT teams to ensure consistent adoption of standards and integration across domains. • Collaborate with OT teams to design and secure infrastructure supporting ICS, DCS, PLCs, SCADA, and plant network environments. • Support secure IT/OT convergence while respecting operational constraints, safety, and uptime requirements. • Ensure OT infrastructure designs align with industry best practices and regulatory expectations for industrial environments. • Embed security-by-design principles into all IT and OT infrastructure architectures. • Define and enforce cybersecurity architecture controls aligned to frameworks such as NIST CSF, NIST 800-53/82, IEC 62443, and Zero Trust principles. • Partner with cybersecurity teams to design controls for network segmentation, identity and access management, monitoring, logging, vulnerability management, and incident response. • Participate in architecture and security risk assessments; identify gaps and drive remediation through design and engineering standards. • Ensure infrastructure solutions comply with corporate policies, cybersecurity standards, and regulatory requirements. • Support audits, risk assessments, and tabletop exercises by providing architectural insight and documentation. • Maintain accurate architecture documentation, diagrams, and decision records. • Serve as a trusted technical advisor to IT, OT, cybersecurity, and business stakeholders. • Provide architectural guidance during project intake, design reviews, and technology evaluations. • Mentor infrastructure engineers and encourage knowledge transfer to reduce key-person dependencies. • Evaluate emerging technologies and recommend improvements to enhance security, resilience, and cost efficiency.




