
TensorWave
Remote Jobs
GPU poor? Contact us for your AI cloud compute needs!
17 Jobs
• Own front-end network architecture - DCI, edge, ingress/egress, and control-plane networks • Architect and operate large edge and service networks • Design scalable Ethernet architectures • Define routing, segmentation, and isolation strategies • Lead hands-on deployment, validation, and troubleshooting in new data centers • Define and maintain reference architectures, standards, and long-term growth models • Own relationships with network carriers and service providers • Work in collaboration with the platform team to design and deliver network solutions for Kubernetes-centric use cases • Partner closely with the Back End Network Principal to define clean interface boundaries between front-end and RDMA back-end fabrics
• Lead all phases of design management for hyperscale and AI data center developments. • Manage architectural and engineering consultants throughout schematic design, design development, construction documents, and permitting. • Coordinate multidisciplinary teams including civil, structural, architectural, mechanical, electrical, plumbing, fire protection, and telecommunications. • Conduct constructability reviews and value engineering exercises to improve project performance while reducing cost and schedule risk. • Partner with Development, Construction, Procurement, and Operations teams to ensure seamless project execution. • Review and approve drawings, specifications, RFIs, submittals, and design changes. • Drive design schedules and monitor consultant deliverables to meet aggressive project milestones. • Ensure compliance with applicable building codes, industry standards, and mission-critical design requirements. • Support procurement of long-lead equipment by validating technical specifications. • Assist in development and maintaining standardized programs, processes, technical standards, workflows, templates, and governance supporting data center design, engineering, construction, commissioning, program management, and operations. • Collaborate with stakeholders to implement best practices that improve quality, consistency, and project execution. • Participate in commissioning planning and operational turnover activities. • Identify project risks and develop mitigation strategies throughout the design lifecycle. • Present project updates and design recommendations to executive leadership and key stakeholders.
Senior Data Center Electrical Engineer
TensorWaveGPU poor? Contact us for your AI cloud compute needs!
• Perform detailed design review and technical validation of high-power electrical distribution systems, including medium voltage switchgear, uninterruptible power supplies (UPS), paralleling switchgear, and generator plants. • Design and model electrical infrastructure to support next-generation AI hardware, focusing on high-density power delivery at the rack level and ensuring system reliability (fault tolerance, selectivity). • Conduct power quality studies, short circuit analysis, and arc flash analysis using industry software (e.g., SKM, ETAP). • Provide subject matter expertise during construction, commissioning (Level 4/5 testing), and operation to resolve complex electrical issues. • Ensure all electrical designs adhere to NFPA 70 (NEC), IEEE, and corporate safety standards.
• Oversee the entire global data center construction portfolio, managing a multi-billion-dollar capital expenditure budget and delivery pipeline. • Develop and execute the program-level construction strategy, including selection and management of general contractors (GCs), EPCM firms, and key vendor partners. • Accountable for ensuring all projects are delivered on-time, within budget, and to the specified quality and technical standards required for AI compute operations. • Establish program-wide construction risk management, quality assurance/quality control (QA/QC) protocols, and project reporting metrics. • Collaborate with Site Acquisition, Design, and Sourcing teams to de-risk projects and resolve complex technical or contractual issues at an executive level.
• Resolve Complex Escalations: Act as the final authority on issues exceeding GOC scope, utilizing code-level debugging and architectural investigation. • Direct Customer Engagement: Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution through active collaboration. • Iterative Problem Solving: Develop diagnostic scripts and workarounds to maintain customer operations while long-term patches are in development. • Drive Root Cause Analysis: Own end-to-end P1 resolution, partnering with TAMs to deliver clear, actionable post-incident analysis. • Bridge to Engineering: Convert recurring customer pain points into evidence-based feature requests, influencing product roadmap to resolve systemic failures. • Build Scalable Knowledge: Document non-obvious platform behaviors and refine GOC runbooks, ensuring institutional knowledge grows with every incident.
Infrastructure Engineer – Storage Platform
TensorWaveGPU poor? Contact us for your AI cloud compute needs!
• Operate and maintain distributed storage platforms, including Ceph (RBD, CephFS, RGW), High-performance NAS platforms (e.g., Weka, VAST Data) • Manage storage lifecycle operations - cluster expansion, upgrades and migrations • Monitor and maintain storage health, including capacity utilization, data distribution and balance, cluster state and recovery operations • Analyze and troubleshoot storage performance across IOPS, throughput, and latency (including tail latency) • Identify and remediate bottlenecks across disk subsystems, network paths (including RDMA where applicable), client access patterns • Support incident response and root cause analysis for storage-related issues • Ensure storage platforms meet performance expectations for GPU and Kubernetes workloads • Operate and support Kubernetes-integrated storage - CSI drivers, StorageClasses, PersistentVolumes / PersistentVolumeClaims • Troubleshoot storage-related issues in Kubernetes environments, including stateful workloads, performance inconsistencies, scheduling and provisioning failures • Execute and improve automation for storage deployment and operations using Ansible, Terraform, Kubernetes manifests / Helm • Contribute to improving monitoring and alerting, operational workflows, runbooks and documentation • Partner with DevOps and Platform Engineering (automation and orchestration), Network Engineering (high-throughput and RDMA networking), Compute / Virtualization teams • Help ensure end-to-end performance across compute, network, and storage layers
DevOps Engineer – Platform Integrations
TensorWaveGPU poor? Contact us for your AI cloud compute needs!
• Design and build integrations between internal platform services (e.g., messaging/pub-sub systems), infrastructure systems (compute, storage, networking), third-party vendor platforms • Develop services and tools that enable automation workflows, system coordination and orchestration, event-driven infrastructure operations • Write production-quality code in Go, Python, Rust (where applicable) • Build APIs, services, and background workers that interact with infrastructure platforms, CI/CD systems, automation frameworks • Ensure code is reliable, observable, maintainable • Integrate software with automation systems such as Ansible, Terraform, CI/CD pipelines (GitHub Actions, ArgoCD) • Enable infrastructure workflows through APIs, event-driven systems, automation hooks • Work closely with DevOps engineers (infrastructure and automation), Development teams (application requirements) • Translate infrastructure capabilities into usable APIs and services • Help teams integrate their systems into platform workflows • Build logging, metrics, and tracing into services • Debug and resolve issues across distributed systems • Ensure integrations are resilient and handle failure scenarios gracefully • Identify gaps in platform integration and automation • Build tooling that reduces manual work and improves system cohesion • Contribute to standards for internal platform development
• Define and own the global data center white space, shell, and critical infrastructure (CI) design standards to support next-generation AI/ML compute hardware, with a focus on power density (e.g., >50kW per rack). • Lead the initial architecture and engineering phases for all purpose-built data center programs, ensuring designs meet hyperscale requirements for reliability (Tier III/IV), efficiency (low PUE), and accelerated deployment. • Direct the selection and integration of advanced cooling topologies, including the strategic adoption of liquid-first and direct-to-chip cooling solutions. • Manage and mentor a team of electrical, mechanical, and architectural design engineers, serving as the final technical authority on all DC engineering decisions. • Drive continuous improvement in design processes, leveraging lessons learned from deployment and operations to reduce time-to-market and construction cost.
• Design and maintain a modular, scalable frontend architecture using TypeScript and React, ensuring consistency across our dashboards • Develop high-performance, reusable component libraries using shadcn/ui to accelerate development and ensure design system integrity • Implement advanced state management and performance patterns to handle complex, real-time GPU telemetry and high-density data visualizations • Establish and enforce frontend standards, including accessibility (WCAG), automated testing, and comprehensive documentation to maintain a high bar for engineering excellence • Partner with backend engineers to create seamless interfaces between our frontend layer and backend services, ensuring reliable data flow and low-latency interactions • Partner closely with UX/UI designers to translate high-fidelity mocks into intuitive user interfaces • Collaborate with cross-functional teams to align frontend strategy with platform goals • Mentor junior engineers through code reviews, pairing sessions, and knowledge sharing to cultivate a culture of continuous learning and high-quality craftsmanship
• Serve as the on-site owner’s representative, managing all phases of a major data center construction project from groundbreaking through final commissioning and handover. • Manage the General Contractor and all associated sub-trades, ensuring strict adherence to the project schedule, technical specifications, and safety plan (EHS). • Own the project budget, tracking all expenditures, change orders, and progress payments to maintain financial control. • Lead daily and weekly project meetings, coordinate the RFI (Request for Information) and submittal process, and resolve field-level issues in partnership with the Design and Engineering teams. • Ensure construction activities comply with corporate quality standards and turnover documentation requirements for the Operations team.
7more opportunities are still waiting for you.Log in now and take your next shot before someone else does.