T

TwelveLabs

Remote Jobs

TwelveLabs is a San Francisco, California based company founded in 2021 that specializes in multimodal, video native, AI for enterprise video understanding. Its

5 open rolesLatest: Apr 1, 2026, 8:00 PM UTCCompany Site
Post Date
Minimum Salary
Experience

5 Jobs

Solutions Engineer

TwelveLabs

TwelveLabs is a San Francisco, California based company founded in 2021 that specializes in multimodal, video native, AI for enterprise video understanding. Its

Solutions Engineer119 days ago

Who we are At TwelveLabs, we are pioneering the development of cutting-edge multimodal foundation models that have the ability to comprehend videos just like humans do. Our models have redefined the standards in video-language modeling, empowering us with more intuitive and far-reaching capabilities, and fundamentally transforming the way we interact with and analyze various forms of media. With $107 million in Seed and Series A funding, our company is backed by top-tier venture capital firms such as NVIDIA's NVentures, NEA, Radical Ventures, and Index Ventures, and prominent AI visionaries and founders such as Fei-Fei Li, Silvio Savarese, Alexandr Wang and more. Headquartered in San Francisco, with an influential APAC presence in Seoul, our global footprint underscores our commitment to driving worldwide innovation. We are a global company that values the uniqueness of each person's journey. It is the differences in our cultural, educational, and life experiences that allow us to constantly challenge the status quo. We are looking for individuals who are motivated by our mission and eager to make an impact as we push the bounds of technology to transform the world. Join us as we revolutionize video understanding and multimodal AI. About the team The Field Engineering team is responsible for ensuring the safe and effective deployment of TwelveLabs' video understanding APIs for developers and enterprises. We act as a trusted advisor and thought partner for our customers, ensuring they maximize value from our multimodal foundation models and products. The Solutions team, part of Field Engineering, is focused on accelerating the sales process for high-impact prospects and customers. They work closely with Sales to articulate technical value, design compelling proof-of-concepts, and demonstrate how our platform solves critical business challenges. About the role We're hiring a Solutions Engineer to help prospects discover and validate the transformative potential of video AI. Acting as the technical expert during the sales cycle, you'll partner with Account Executives to qualify opportunities, scope use cases, design and execute proof-of-concepts, and demonstrate the business value of our multimodal foundation models. You'll provide technical leadership throughout the sales process while collaborating with Sales, Product, Partnerships and Engineering teams. This is a highly self-directed and creative role where you'll need to thrive in an unstructured, rapidly expanding and evolving environment. This role is remote. In this role, you will Partner with Account Executives to drive technical discovery, understand prospect challenges, and articulate how TwelveLabs' video understanding APIs solve their business problems. Design and execute compelling proof-of-concepts that demonstrate measurable business value and technical feasibility, helping prospects envision the art of the possible. Deliver technical presentations, product demonstrations, and workshops to diverse audiences from developers to C-suite executives, showcasing the capabilities of our platform. Serve as the technical voice in sales cycles, addressing questions around architecture, security, compliance, integration patterns, and scalability. Gather technical requirements and competitive intelligence from prospects to inform Sales strategy and represent market needs to Product and Engineering teams. Create reusable demo applications, solution architectures, and technical collateral that accelerate the sales process and showcase best practice You might be a good fit if you have Have 5+ years of experience in a technical, customer-facing role such as Sales Engineer, Solutions Engineer, or Solutions Architect with a focus on pre-sales activities. Demonstrate a strong understanding of and experience working with LLMs, multimodal AI, REST APIs, Python, JavaScript, and modern application development, with a deep understanding of video workflows and video workloads. Have experience designing and executing proof-of-concepts that effectively demonstrate ROI and technical feasibility to accelerate sales cycles. Are an exceptional communicator and presenter who can compellingly articulate technical value to business stakeholders and build credibility with technical audiences. Operate with high horsepower, manage multiple opportunities simultaneously, and thrive in a fast-paced sales environment with evolving priorities. Bring a collaborative, competitive attitude with natural curiosity about customer challenges and eagerness to win deals alongside Sales teams. Even if there are a few checkboxes that aren't ticked through your prior experience, we still encourage you to apply! If you are a 0-to-1 achiever, a ferocious learner, and a kind and fun team player who motivates others, you will find a home at TwelveLabs. We welcome applicants from all walks of life and are committed to equal-opportunity employment. We cherish and celebrate diversity not just because it is the right thing to do, but because it makes our company much stronger.

United Kingdom

Staff Software Engineer, Platform Integrations

TwelveLabs

TwelveLabs is a San Francisco, California based company founded in 2021 that specializes in multimodal, video native, AI for enterprise video understanding. Its

Software Engineer124 days ago

Who we are At Twelve Labs, we are pioneering the development of cutting-edge multimodal foundation models that have the ability to comprehend videos just like humans do. Our models have redefined the standards in video-language modeling, empowering us with more intuitive and far-reaching capabilities, and fundamentally transforming the way we interact with and analyze various forms of media. With a remarkable $107 million in Seed and Series A funding, our company is backed by top-tier venture capital firms such as NVIDIA’s NVentures, NEA, Radical Ventures, and Index Ventures, and prominent AI visionaries and founders such as Fei-Fei Li, Silvio Savarese, Alexandr Wang and more. Headquartered in San Francisco, with an influential APAC presence in Seoul, our global footprint underscores our commitment to driving worldwide innovation. We are a global company that values the uniqueness of each person’s journey. It is the differences in our cultural, educational, and life experiences that allow us to constantly challenge the status quo. We are looking for individuals who are motivated by our mission and eager to make an impact as we push the bounds of technology to transform the world. Join us as we revolutionize video understanding and multimodal AI. About the Role You'll own the infrastructure and integration layer that makes TwelveLabs models available on partner platforms. This is everything outside the model itself: how model containers are packaged, validated, and deployed; how API surfaces are designed and maintained per platform; how requests are routed; and how we ensure production reliability across fundamentally different cloud environments. You'll work closely alongside our Science, Product, and ML Engineering teams to align the model and product roadmap for effective platform integrations. Your domain is the external model orchestration — you need to understand how model components function (to make good integration decisions), but you won't be optimizing the models themselves. Your work accelerates the ability to reliably ship new model versions and features to users across all platforms. Candidates must be able to travel up to 10% of the time annually to attend conferences, off-site meetings, and other business-related events as required by the role. This role may require participation in on-site interviews and/or completion of in-person onboarding processes. In this role, you will - Design and build infrastructure that deploys TwelveLabs models across multiple cloud and data platforms, accounting for differences in compute hardware, networking, APIs, and operational models - Own direct integrations into partner products — implementing the orchestration, data flow, and API surfaces that connect TwelveLabs models to partner-side functionality - Design and evolve CI/CD automation systems — including validation and deployment pipelines that reliably ship new model versions across platforms without regressions - Design interfaces and tooling abstractions across platforms that enable consistent deployment, reduce per-platform complexity, and scale as we add new partners - Implement API-level features and changes that require understanding model component behavior — routing, request handling, response formatting — without modifying model internals - Contribute to capacity planning and autoscaling strategies that dynamically match supply with demand across platform deployments - Analyze observability data across platforms to identify performance bottlenecks, cost anomalies, and regressions — and drive remediation based on production workloads - Collaborate with platform partner engineering teams to resolve operational issues, align on API contracts, and stand up end-to-end serving on new platforms You may be a good fit if you have - Significant software engineering experience building and operating mission-critical backend systems at scale - Experience building or operating services on at least one major cloud platform (AWS, GCP, or Azure), with exposure to Kubernetes, infrastructure as code, or container orchestration - Strong interest in ML inference — you want to understand how models work, even if your primary contribution is the infrastructure around them - Ability to design highly observable systems that operate reliably at scale across multiple environments - Autonomy and ownership — you take problems end to end with a bias toward high-impact work Preferred Qualifications - Direct experience working with cloud provider partner teams to scale infrastructure or products across multiple platforms — navigating differences in networking, security, billing, and managed service offerings - Background building platform-agnostic tooling or abstraction layers that work across cloud providers - Hands-on experience with capacity management, cost optimization, or resource planning at scale across heterogeneous environments - Familiarity with ML inference optimization, batching, caching, and serving strategies - Experience with ML infrastructure including GPUs, TPUs, Trainium, or other AI accelerators - Background designing CI/CD systems that automate deployment and validation across cloud environments - Proficiency in Python or Go Benefits and Perks 🤝 An open and inclusive culture and work environment. 🚀 Work closely with a collaborative, mission-driven team on cutting-edge AI technology. 🏥 Full health, dental, and vision benefits ✈️ Extremely flexible PTO and parental leave policy. Office closed the week of Christmas and New Years. 🛂 VISA support where applicable

United States

Principal Software Engineer, Video Engineering

TwelveLabs

TwelveLabs is a San Francisco, California based company founded in 2021 that specializes in multimodal, video native, AI for enterprise video understanding. Its

Software Engineer124 days ago

Who we are At Twelve Labs, we are pioneering the development of cutting-edge multimodal foundation models that have the ability to comprehend videos just like humans do. Our models have redefined the standards in video-language modeling, empowering us with more intuitive and far-reaching capabilities, and fundamentally transforming the way we interact with and analyze various forms of media. With a remarkable $107 million in Seed and Series A funding, our company is backed by top-tier venture capital firms such as NVIDIA’s NVentures, NEA, Radical Ventures, and Index Ventures, and prominent AI visionaries and founders such as Fei-Fei Li, Silvio Savarese, Alexandr Wang and more. Headquartered in San Francisco, with an influential APAC presence in Seoul, our global footprint underscores our commitment to driving worldwide innovation. We are a global company that values the uniqueness of each person’s journey. It is the differences in our cultural, educational, and life experiences that allow us to constantly challenge the status quo. We are looking for individuals who are motivated by our mission and eager to make an impact as we push the bounds of technology to transform the world. Join us as we revolutionize video understanding and multimodal AI. About The Role Most video engineering roles at Netflix, YouTube, or Twitch optimize for human playback — better compression, lower bitrate, smoother streaming. At Twelve Labs, video is processed for machine understanding. The tradeoffs are fundamentally different: we optimize for AI model performance, not just perceptual quality. This is a rare opportunity to define how video is engineered for intelligence — not just delivery. As the Principal Software Engineer, Video Engineering, you will own the architecture and implementation of Twelve Labs' video processing pipelines — from byte ingestion through decode, chunking, storage, and playback — ensuring it is fast, cost-efficient, and purpose-built for AI-native video intelligence at scale. You will be the internal subject matter expert on all things related to video engineering. In This Role You Will: - Own the video pipeline end-to-end: Architect and implement ingestion → decode → chunking → storage → retrieval → playback, across batch and streaming modes based on AI/ML workflows or media application workflows. - Deep codec & decode mastery: Drive decisions on decode strategies (hardware vs. software, GPU-accelerated pipelines), container format handling (fMP4, CMAF, MKV, TS), and codec support (H.264, H.265, VP9, AV1) with pragmatic cost/quality tradeoffs. - Semantic & heuristic chunking: Work with our ML Research Scientists to design and implement intelligent video segmentation that goes beyond fixed-interval splitting — scene boundary detection, shot change analysis, content-aware chunking that optimizes downstream AI model performance. - Streaming ingestion: Architect low-latency streaming pipelines (HLS, DASH, LL-HLS, WebRTC ingest) that process video in near-real-time, including streaming decode and incremental chunking. - Video storage architecture: Design storage tiers and retrieval patterns optimized for AI workloads — balancing hot/warm/cold access, frame-level random access, and cost at petabyte scale. - Playback & delivery: Ensure video can be served back to users with accurate temporal navigation, supporting time-coded references from AI analysis results. - FFmpeg & media toolchain expertise: Be the internal authority on FFmpeg, libav, and related tooling. Build and maintain custom processing pipelines, filters, and integrations. - Cost engineering: Quantify and optimize cost-per-hour-of-video-processed. Drive decode efficiency through hardware acceleration (NVDEC, VA-API), pipeline parallelism, and intelligent resource allocation. - Cross-team technical leadership: Partner with ML teams on how video is preprocessed for model consumption, with platform teams on infrastructure, and with product on customer-facing media capabilities. - Standards & best practices: Establish video engineering standards, author reference implementations, and mentor engineers across teams on media fundamentals. You May Be A Good Fit If You Have: - 12+ years in software engineering with 7+ years focused on video/media engineering in production systems processing video at scale. - Deep FFmpeg expertise: Not just CLI usage — understanding of libavcodec, libavformat, filter graphs, custom demuxers/decoders, and performance tuning. - Codec internals knowledge: H.264/H.265 bitstream structure, AV1 adoption tradeoffs, hardware decode paths, quality metrics (VMAF, SSIM, PSNR). - Streaming protocol fluency: HLS, DASH, LL-HLS, WebRTC. Experience with live/real-time ingest pipelines. - Systems engineering depth: Comfortable in C/C++, Rust, or Go for performance-critical media code; Python for pipeline orchestration. Can reason about memory layout, SIMD, GPU pipelines. - Storage & retrieval at scale: Experience designing video storage systems — object stores, frame-indexed access patterns, tiered storage strategies. - Content-aware processing: Experience with scene detection, shot boundary analysis, temporal segmentation, or perceptual quality optimization. - Production instincts: Incident response, observability for media pipelines, debugging decode failures at scale, handling format edge cases gracefully. - AI/ML integration experience (strongly preferred): Worked with teams consuming video frames for model training/inference. Understands how preprocessing decisions (resolution, frame rate, chunking strategy) impact model quality. Qualified Candidates May Also Have: - Made major contributions to FFmpeg, GStreamer, or open-source media projects. - Deep familiarity with GPU-accelerated video processing (ex. NVDEC/NVENC). - Experience running media pipelines in constrained environments such as on-prem or edge settings. Candidates must be able to travel up to 10% of the time annually to attend conferences, off-site meetings, and other business-related events as required by the role. This role may require participation in on-site interviews and/or completion of in-person onboarding processes. Benefits and Perks 🤝 An open and inclusive culture and work environment. 🚀 Work closely with a collaborative, mission-driven team on cutting-edge AI technology. 🏥 Full health, dental, and vision benefits ✈️ Extremely flexible PTO and parental leave policy. Office closed the week of Christmas and New Years. 🛂 VISA support where applicable

United States

Research Scientist, Public Sector

TwelveLabs

TwelveLabs is a San Francisco, California based company founded in 2021 that specializes in multimodal, video native, AI for enterprise video understanding. Its

Research Scientist128 days ago

Who we are At Twelve Labs, we are pioneering the development of cutting-edge multimodal foundation models that have the ability to comprehend videos just like humans do. Our models have redefined the standards in video-language modeling, empowering us with more intuitive and far-reaching capabilities, and fundamentally transforming the way we interact with and analyze various forms of media. With a remarkable $107 million in Seed and Series A funding, our company is backed by top-tier venture capital firms such as NVIDIA’s NVentures, NEA, Radical Ventures, and Index Ventures, and prominent AI visionaries and founders such as Fei-Fei Li, Silvio Savarese, Alexandr Wang and more. Headquartered in San Francisco, with an influential APAC presence in Seoul, our global footprint underscores our commitment to driving worldwide innovation. We are a global company that values the uniqueness of each person’s journey. It is the differences in our cultural, educational, and life experiences that allow us to constantly challenge the status quo. We are looking for individuals who are motivated by our mission and eager to make an impact as we push the bounds of technology to transform the world. Join us as we revolutionize video understanding and multimodal AI. About the Role As a Research Scientist on the Public Sector team, you will adapt and deploy TwelveLabs' video AI capabilities - including our multimodal foundation model - for mission critical government applications. This role focuses on applying our video intelligence technology to classified and government-specific use cases, operating within on-premises, GovCloud, and air-gapped environments. You will be the dedicated research scientist for the Public Sector team, bridging TwelveLabs' cutting-edge multimodal AI research and the unique requirements of U.S. federal, defense, and intelligence community customers. In this role, you will: - Adapt TwelveLabs' video understanding and multimodal models for government-specific use cases (defense, intelligence analysis, federal records management) - Support deployment of models in air-gapped, GovCloud, and on-premises environments with strict compliance requirements - Develop evaluation methodologies and data strategies specific to public sector applications - Ensure all model deployments meet government security standards (FedRAMP, DoD SRG, FIPS) - Work closely with Solutions Engineering to translate customer requirements into technical implementations You may be a good fit if you have: - Strong research experience in areas such as: multimodal understanding, large language models, representation learning, computer vision, or domain adaptation - Experience operationalizing AI/ML in DoD or Intelligence Community - Proficiency in Python and PyTorch - Experience deploying ML models in production environments - Understanding of air-gapped / disconnected environment constraints - Active Top Secret clearance or ability to obtain - PhD or Master's in Computer Science, Mathematics, or related field Strong candidates may also have: - Active TS/SCI clearance - Experience operationalizing AI/ML in DoD or Intelligence Community - Experience with government deployment requirements (FedRAMP, FIPS, air-gapped networks) - Background in video understanding or video-language models - Publications in top conferences (CVPR, NeurIPS, etc.) Candidates must be able to travel up to 10% of the time annually to attend conferences, off-site meetings, and other business-related events as required by the role. This role may require participation in on-site interviews and/or completion of in-person onboarding processes. Benefits and Perks 🤝 An open and inclusive culture and work environment. 🚀 Work closely with a collaborative, mission-driven team on cutting-edge AI technology. 🏥 Full health, dental, and vision benefits ✈️ Extremely flexible PTO and parental leave policy. Office closed the week of Christmas and New Years. 🛂 VISA support where applicable

United States

Staff Product Designer

TwelveLabs

TwelveLabs is a San Francisco, California based company founded in 2021 that specializes in multimodal, video native, AI for enterprise video understanding. Its

Product Designer148 days ago

Who we are At Twelve Labs, we are pioneering the development of cutting-edge multimodal foundation models that have the ability to comprehend videos just like humans do. Our models have redefined the standards in video-language modeling, empowering us with more intuitive and far-reaching capabilities, and fundamentally transforming the way we interact with and analyze various forms of media. With a remarkable $107 million in Seed and Series A funding, our company is backed by top-tier venture capital firms such as NVIDIA’s NVentures, NEA, Radical Ventures, and Index Ventures, and prominent AI visionaries and founders such as Fei-Fei Li, Silvio Savarese, Alexandr Wang and more. Headquartered in San Francisco, with an influential APAC presence in Seoul, our global footprint underscores our commitment to driving worldwide innovation. We are a global company that values the uniqueness of each person’s journey. It is the differences in our cultural, educational, and life experiences that allow us to constantly challenge the status quo. We are looking for individuals who are motivated by our mission and eager to make an impact as we push the bounds of technology to transform the world. Join us as we revolutionize video understanding and multimodal AI. About the Role TwelveLabs is building the platform that every intelligent video workflow will run on. Our products don't analyze video frame-by-frame, instead they see, hear, and reason across it the way a human does. We've built the most powerful video understanding engine in the world. Now we're designing the experience layer for everyone to build on. This role defines design across the entire TwelveLabs product surface including applications, developer platform, APIs, and the emerging products connecting them. The design challenges are genuinely unsolved: How should someone interact with AI that operates on entire corpuses of video? What does a creative tool look like when the AI already knows where the story beats are? How do you make probabilistic output feel trustworthy? The patterns you set here won't just define TwelveLabs, they'll define how the world interacts with video intelligence. This role will make the concept of video chat obsolete. What you'll do - Help define the design language and interaction model across every TwelveLabs product surface: applications, platform, APIs, brand, launch materials - Invent interaction and design patterns for a new paradigms across all scales from single frames to entire video corpuses - Prototype relentlessly. Test with developers in production, and enterprise teams working across millions of hours of footage - Build the design system by shipping You might be a fit if - You've shipped products people love, not products that survived - You’ve found a way to maximize generative AI in your work, it isn’t replacing you it’s elevating you - You move at a pace that surprises engineers and burn more tokens than they do - You understand video and other types of multimodal content at their DNA - You've designed for both creators and developers. Consumer polish and developer experience rigor - You've worked with AI products or you're obsessed with the design challenges of non-deterministic systems - Strong taste, strong opinions, zero preciousness Design Philosophy: We're building the platform layer for video intelligence. There are no established UX conventions for this. Every pattern you ship becomes the reference for us, for our customers, and eventually for the industry. No handoffs. No design-by-committee. No enterprise software wearing a nice font. We're designing for a world where the AI has already helped with the hard part and the human's job is to direct, refine, and decide. First-principles thinking, not copying what existing tools did. Speed is non-negotiable. Conviction over consensus. Product over process. We are a global company that values the uniqueness of each person’s journey. It is the differences in our cultural, educational, and life experiences that allow us to constantly challenge the status quo. We are looking for individuals who are motivated by our mission and eager to make an impact as we push the bounds of technology to transform the world. Join us as we revolutionize video understanding and multimodal AI. Benefits and Perks 🤝 An open and inclusive culture and work environment. 🚀 Work closely with a collaborative, mission-driven team on cutting-edge AI technology. 🏥 Full health, dental, and vision benefits ✈️ Extremely flexible PTO and parental leave policy. Office closed the week of Christmas and New Years. 🛂 VISA support where applicable

United States