
Runpod
Remote Jobs
AI Developer Cloud
51 Jobs
• Assist in validating new hardware, ensuring partner deployments meet Runpod’s specifications for distributed AI/ML workloads. • Monitor fleet health to identify performance degradation. You will help audit downtime and provide the technical data needed to protect customer SLAs. • We operate with an AI-first mindset, powering our operations with the technology we host. You will work with LLMs and AI agents to help automate network triage and generate dynamic runbooks for our fleet. • Coordinate technical incident communications with clear updates, acting as a steady hand that translates outages into actionable resolutions. • Support the growth of our infrastructure partners.
• Assist in building, maintaining, and enhancing user interfaces, APIs, and other core components of our platform using Python, Typescript, and Go Lang. • Work with our open-source SDK to ensure it remains intuitive and effective for developers integrating with Runpod’s platform. • Support cross-functional teams by developing features and addressing bugs that enhance the user experience for AI engineers and researchers. • Write clean, maintainable code with an emphasis on performance and scalability, following best practices for testing and continuous integration. • Participate in code reviews, mentorship opportunities, and continuous learning to deepen your expertise in full-stack development.
• Design, develop, and maintain cloud infrastructure software, primarily in Go. • Collaborate with managers and other engineers to define and implement product requirements. • Troubleshoot and optimize existing code to improve performance and reliability. • Participate in code reviews and contribute to the team's technical standards. • Contribute to architectural discussions and decisions. • Stay up-to-date with industry trends and emerging technologies.
• Own and execute a multi-quarter product strategy for a high-impact area of Runpod’s AI infrastructure platform, directly tied to revenue growth, adoption, retention, gross margin, customer outcomes, and platform expansion. • Define the target market, customer segments, buyer personas, user personas, and ideal customer profile using customer research, product usage data, sales insights, competitive analysis, and market trends. • Build business cases for major product investments, including market opportunity, customer pain points, willingness to pay, competitive dynamics, adoption barriers, risks, and expected ROI. • Develop and manage a KPI-driven roadmap where every major initiative has a clear customer problem, business objective, success metric, leading indicator, and measurable impact. • Define, track, and report on key product and business metrics, including activation, conversion, usage growth, revenue contribution, retention, expansion, churn, compute utilization, gross margin, time-to-value, NPS, and feature engagement. • Use quantitative and qualitative data to guide product decisions, prioritize roadmap tradeoffs, evaluate launch performance, and continuously refine strategy. • Conduct frequent customer discovery with prospects, customers, partners, and internal go-to-market teams to identify scalable market opportunities and validate product direction. • Partner with Engineering, Design, and Infrastructure to deliver reliable, scalable, and intuitive platform capabilities that improve customer outcomes and support business growth. • Collaborate with Sales, Marketing, Customer Success, Finance, Partnerships, and Legal to drive go-to-market readiness, including positioning, pricing and packaging input, launch planning, enablement, onboarding, and expansion motions. • Communicate product strategy, business impact, roadmap progress, competitive insights, key risks, and tradeoffs clearly to executives, strategic customers, and cross-functional stakeholders.
• Own Distributed Storage Architecture: Define, evolve, and operate Runpod’s global storage platforms, supporting training, inference, checkpointing, and dataset access at scale. • Build the Storage Engineering Team: Manage and grow a team of storage and systems engineers. Set clear ownership, technical direction, and operational standards across regions. • High-Performance Shared Filesystems: Design and operate large-scale SAN and NFS deployments, including performance-sensitive shared storage for GPU clusters. • Advanced Filesystems & Platforms: Lead deployments and operations of VAST Data and experience with Lustre or similar parallel filesystems used in HPC and AI environments. • End-to-End Performance Ownership: Drive performance optimization from NAND and NVMe media through controllers, networking, and client access patterns. • Next-Generation Storage Technologies: Evaluate and deploy cutting-edge capabilities such as NFS over RDMA, GPU Direct Storage (GDS), and low-latency data paths for accelerated workloads. • Reliability & Scale: Establish best practices for replication, data tiering, data protection, failure recovery, capacity planning, and lifecycle management. • Automation & Observability: Build automation for provisioning, expansion, upgrades, and monitoring. Ensure deep observability into throughput, latency, and error characteristics. • Cross-Functional Collaboration: Partner with Datacenter Networking, GPU Platform, SRE, and Product teams to ensure storage systems meet evolving workload and customer needs. • Vendor & Partner Management: Own technical relationships with storage vendors, hardware partners, and colocation providers; drive roadmap alignment and issue resolution.
• Serve as a strategic partner to the COO and executive team on people strategy and organizational design • Support executives in scaling teams, refining organizational structures, and strengthening leadership capabilities • Lead change management to ensure smooth transitions and transformation during growth • Align HR initiatives directly with business objectives, ensuring people strategy drives business results • Design company-wide people programs and frameworks (performance management, compensation, career development) • Own and optimize people operations systems, processes, and technology stack (HRIS, benefits platforms, compliance tools) • Collaborate with the Director of Talent Acquisition, who owns recruiting execution, to align hiring strategy, headcount planning, and operational scalability • Build and manage operational infrastructure including benefits strategy, vendor relationships, and compliance • Manage people operations budget including headcount, benefits, tools, and programs • Ensure compliance with federal, state, and local employment laws across all locations • Design compensation and rewards structures that are competitive and equitable • Administer performance management systems that encourage accountability and excellence • Lead, mentor, and support a small but impactful People team • Provide clear direction while empowering team members to own their areas of expertise • Set strategic direction for how the People team operates and delivers value • Build additional people operations capability as the company scales • Set strategic direction for culture, engagement, and employee experience • Partner with Director, HRBP to ensure consistent implementation across all business functions • Foster an environment where employees feel engaged, supported, and motivated to grow
• Marketing operations strategy, roadmap, and day-to-day execution across HubSpot, Salesforce, Webflow, analytics, attribution, automation, and reporting systems. • Campaign operations workflows for launches, events, partner programs, lifecycle campaigns, paid programs, and content distribution. • Funnel reporting across self-serve, PQL, MQL, sales handoff, pipeline, and revenue influence. • Lead routing, scoring, segmentation, lifecycle automation, source attribution, and data quality processes. • Weekly and monthly marketing operating reviews, including dashboards, performance readouts, and decision-ready reporting for leadership. • Cross-functional operating rhythm with RevOps, Sales, Finance, Product, and Data so marketing metrics are trusted and actionable. • QA processes for campaign tracking, forms, landing pages, UTMs, CRM fields, lifecycle workflows, and handoff rules. • Documentation for definitions, workflows, attribution logic, launch checklists, and source-of-truth reporting.
• Serve as the main point of contact for a portfolio of existing accounts. • Monitor account health metrics (usage, billing, feature adoption, churn risk) and proactively intervene. • Drive renewals, upsells, and cross-sells by identifying expansion opportunities (e.g. more GPU capacity, new feature modules, premium support plans). • Conduct business reviews (QBR/MBR) to update clients on ROI, roadmap, usage insights, and new offerings. • Liaise with product, engineering, and support teams to resolve customer issues, advocate for feature enhancements, and remove blockers. • Build account plans, forecasts, and growth strategies for each client. • Track key account metrics (e.g. churn rate, net revenue retention, expansion revenue, customer satisfaction) and report them to leadership. • Handle contract renewals, negotiation, and escalations. • Maintain accurate data in CRM (deal pipeline, account notes, forecasts) • Occasionally assist in onboarding or transitions of new clients within your portfolio. • Periodic travel to visit top customers or represent Runpod at key industry events.
• Lead threat modeling, architecture reviews, and code reviews for our web applications, APIs, and microservices. • Actively develop and commit code to fix security flaws in our Python, Go, or JavaScript/TypeScript codebases alongside the engineering team. • Implement, tune, and manage security testing tools (SAST, DAST, SCA) within our CI/CD pipelines to catch vulnerabilities early in the SDLC. • Configure and manage application-layer security controls, including Web Application Firewalls (WAF), bot protection, and API gateways. • Provide security guidance, secure coding training, and standard operating procedures to development teams. • Collaborate with operations to ensure product-level adherence to relevant frameworks (e.g., SOC 2, ISO 27001, GDPR) and participate in bug bounty triage.
• Design and implement robust workload and network isolation architectures for RunPod's multitenant GPU bare-metal and virtualized environments. • Harden Linux kernel configurations, container runtimes (e.g., Docker, containerd), and orchestration layers (e.g., Kubernetes) against breakouts and privilege escalation. • Conduct deep-dive security assessments and penetration testing specifically targeting our hypervisor, network stack, and hardware interfaces. • Write code (primarily C, Go, or Rust) to implement custom security controls, telemetry, and fixes at the OS and infrastructure level. • Evaluate and mitigate security considerations specific to GPU architecture, PCIe pass-through, and shared memory spaces. • Serve as the technical escalation point for infrastructure-level security incidents, developing forensic capabilities for ephemeral container environments.
41more opportunities are still waiting for you.Log in now and take your next shot before someone else does.