Cantina logo
Cantina

Building the first social AI platform

Machine Learning Engineer, Speech – Joint Audio-Video Modeling

Machine Learning EngineerMachine Learning EngineerFull TimeRemoteSeniorTeam 51-200Since founded by Sean ParkerH1B SponsorCompany SiteLinkedIn

Location

Europe

Posted

6 days ago

Salary

$200K - $220K / year

Seniority

Senior

Bachelor DegreeExperience acceptedEnglishCloudNode.jsPyTorch

Job Description

Machine Learning Engineer, Speech – Joint Audio-Video Modeling

Cantina

• Audio Representations: Design, train, and improve the audio VAEs, neural codecs, and vocoders our generative models sit on top of latent design, reconstruction and perceptual objectives, compression-vs-fidelity tradeoffs. • Model Building: Architect, implement, pre-train, fine-tune, and post-train/alignment (e.g., GRPO/DPO) diffusion and flow-matching transformers for large-scale audio and video generation. • Joint Audio-Video Modeling: Design the audio conditioning and cross-modal alignment inside joint AV models, audio latents alongside video latents, reference-audio and multi-speaker conditioning, multi shot generation audio/video modeling. • Experimental Design: Design, run, and analyze scientific experiments to advance our understanding of the models. • Data Ownership: Define data requirements and collaborate on acquisition, curation, AV-sync and quality filtering, annotation quality, and synthetic data strategies for paired audio-video and speech corpora. • Rigorous Evaluation: Design automated objective/subjective evaluations audio fidelity and intelligibility metrics, AV-sync, listening and viewing tests, robustness & bias checks, and red-team studies. • Inference Efficiency: Drive distillation, step-count reduction, quantization, and kernel/memory optimization to meet interactive latency and cost targets. • Pipeline Delivery: Harden the training → evaluation → inference pipeline; profile latency, memory, and cost; and meet production SLAs with robust monitoring and rollback. • GPU Scaling: Partner with infrastructure to run distributed training/inference on cloud fleets and productionize models with reliability and observability. • Project Leadership: Independently lead small research projects while collaborating on larger team initiatives, including cross-team work with video generation. • Tool Development: Develop and improve dev tooling to enhance team productivity. • Safety & Responsibility: Contribute to safety/consent guardrails, watermarking, and misuse/abuse mitigation for responsible voice and likeness technology.

Job Requirements

  • Exceptional research/development experience with large-scale audio models (>8B parameters, >500k hours of data).
  • Deep hands-on experience with diffusion and/or flow-matching transformers, including practical knowledge of samplers, schedules, conditioning mechanisms, and distillation.
  • Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders latent/tokenizer design, reconstruction and perceptual objectives, adversarial training.
  • Strong experience with multi-node, multi-GPU distributed training (FSDP/DeepSpeed or equivalent).
  • Strong software engineering skills with a proven track record of building complex systems.
  • Strong with PyTorch and performance work (profiling, CUDA/Triton/C++ as needed) and writing reliable production-quality code.
  • Shipped large-scale speech/audio or multimodal generative models to production.
  • Background in working with large-scale ML data, and the ability to iterate on data and triangulate quality using both subjective and objective signals.
  • Experience with voice cloning, speech control/steerability, or expressive speech generation.
  • Notable publications and/or open-source contributions in speech/audio/ML.
  • Strongly preferred:
  • Experience with multimodal audio-video modeling: joint AV generation of multi-shot, multi-speaker scenes with dialogue, music, and sound design generated jointly with video, and the cross-modal alignment that keeps them in sync.
  • Experience with video generation: video diffusion/flow-matching transformers, video VAEs, conditioned and multi-shot generation, building data pipelines for video models.
  • Streaming or real-time generation, causal distillation (e.g., Self Forcing / Self Forcing++).

Benefits

  • Competitive salary and generous company equity
  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
  • 42 days of paid time off, including:
  • 15 PTO days
  • 10 sick days
  • 15 company holidays
  • 2 floating holidays
  • Generous parental leave & fertility support
  • 401(k) retirement savings plan
  • Lifestyle spending account – $500/month to use however you’d like
  • Complimentary lunch and snacks for in-office employees
  • One Medical membership, and more!

Related Job Pages

More Machine Learning Engineer Jobs

Reddit, Inc logo

Machine Learning Manager, Feed Relevance – Retrieval

Reddit, Inc

Reddit is an online platform utilized by thousands of communities to connect and converse about a wide variety of topics, including TV and movie fan theories, s

• Define the technical vision and long-term roadmap for Feed Retrieval, aligning large-scale recommender-system investments with Reddit’s product, ecosystem, and business objectives. • Translate broad Feed Relevance goals into a focused team roadmap, making clear prioritization tradeoffs across model quality, inventory expansion, experimentation velocity, infrastructure cost, and operational reliability. • Coach and support the development of your team, constantly seeking opportunities to grow their skills and impact. • Oversee the design, development, and optimization of retrieval systems that source relevant, diverse, fresh, and high-quality candidates for personalized feed experiences. • Establish strong measurement, experimentation, and debugging practices so the team can understand retrieval quality, candidate coverage, source incrementality, and downstream impact. • Collaborate with ML platform, infrastructure, ranking, safety, and product teams to build scalable, low-latency retrieval systems that can support the next generation of AI-powered recommendations. • Maintain high standards for system performance, reliability, latency, cost efficiency, and responsible recommendation practices. • Work with cross-functional partners from across the company to identify key areas of opportunity, set expectations, and communicate your team’s work. • Partner with our incredible recruiting team to attract, interview, and hire diverse and talented machine learning engineers, growing a world-class team.

United States
$253.3K - $354.6K / year
Reddit, Inc logo

Machine Learning Manager, Feed Ecosystems

Reddit, Inc

Reddit is an online platform utilized by thousands of communities to connect and converse about a wide variety of topics, including TV and movie fan theories, s

• Define the technical vision and long-term roadmap for Feed Ecosystems. • Coach and support the development of your team. • Work closely with product, design, data science, safety, community, ads, and platform partners. • Oversee the design, development, and optimization of ML systems. • Help define and operationalize signals for subjective and objective quality. • Collaborate with platform and infrastructure teams to build scalable AI-powered systems. • Maintain high standards for system performance, reliability, efficiency, and responsible AI practices. • Partner with our recruiting team to attract, interview, and hire diverse and talented machine learning engineers.

United States
$253.3K - $354.6K / year

Role Description Have you earned your OSCP, OSEP/OSED/OSWE, and gained offensive experience before and/or since then? Are you excited at the opportunity to contribute to the growth and education for the current and next generation of cybersecurity professionals? If the idea of building lab environments to provide hands-on experience for individuals to learn and upskill, then this might be the right role for you! Duties and Responsibilities - Content design and build - Researches and identifies topical, relevant attack vectors suitable for inclusion in OffSec labs, exams, and/or learning content. - Design and build VMs and multi-host environments from vectors across Linux and Windows, including Active Directory and Cloud attack paths. - Designs realistic scenarios and environment narratives that make each attack path plausible and discoverable through enumeration. - Ensures the range of vectors in each environment is appropriate to its stated difficulty level and learning objectives. - Maintains variety across the catalogue so that content remains distinct as it rotates through active use. - Quality, integrity, and reproducibility - Review labs for unintended solution paths and confirms each is solvable by the intended route within its expected time frame. - Validates that machines are deterministic and reproducible in an isolated environment, including reliable reset and revert behaviour. - Deliberate selection and use of software and OS versions, monitoring deployed components so that content behaves consistently over its lifetime. - Documents the steps required to complete each machine, with difficulty assessed at a level of detail that supports competency mapping and certification review. - Collaboration - Coordinates with the relevant team(s) to deploy, update, and retire content. - Coordinates with testing to ensure full coverage of both newly created and updated content. - Coordinates with the Lab team Manager and Content Architect on vector selection and exploit implementation within courses and learning paths. - Coordinates and makes recommendations to keep content additions consistent across products. - Regularly communicates content additions, changes, and retirements to relevant stakeholders. - Senior scope - Acts as a technical reviewer for lab concepts and builds produced by other Vulnerable Machine Engineers and community contributors, providing structured, actionable feedback. - Support and mentor other engineers on build standards, documentation quality, and content review practice. - Contributes across OffSec's lab and exam portfolio as priorities require, rather than to a single product. - Automation and tooling - Creates and maintains automation that makes lab creation timely, consistent, and repeatable. - Evaluates new and emerging technologies to make recommendations on their introduction into the lab estate. - Creates and maintains documentation for every lab, to a standard that allows any team member to rebuild it from the documentation alone, without assistance. Qualifications - OSCP minimum, OSCE³ preferred (or at least one OffSec 300-level cert required). - A Bachelor’s Degree in Systems Engineering, Computer Science, Information Systems, and/or Information Assurance from an accredited institution/related specialized field, or equivalent practical experience. - Demonstrable Active Directory and Cloud attack-path experience. - Proficiency with infrastructure-as-code and configuration management and version control. - Scripting proficiency (Python, PowerShell, Bash). - At least five or more years of experience as a Systems Engineer or in an Info-Sec related position, with experience reviewing or approving others' technical security content. - Experience with Penetration Testing. - Experience with Bug Bounty Hunting. - Experience as a Systems Administrator with a variety of Operating Systems: - MS Windows - MS Windows Based Server Systems - POSIX Server Systems - Linux and Mac Desktop Environments - Strong written and verbal communication skills with an ability to present technical ideas clearly to both technical and non-technical audiences. Work Location and Hours This role is a full-time salaried position. It is a fully remote position. Work hours for this position are flexible and will be performed from a home office. Direct Reports This position has no direct reports. However, the expectations of this role are to provide technical leadership. EEO OffSec provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. This policy applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation and training.

Worldwide
Upstart logo

Senior Applied ML Engineer

Upstart

Our mission is to enable effortless credit based on true risk.

Full TimeRemoteTeam 1,001-5,000H1B Sponsor

• Design and build user-facing ML features that harness LLMs and generative AI to unlock new product capabilities • Partner with product, design, and ML research to prototype and deliver high-impact, ML-powered experiences • Own the technical architecture and implementation strategy for applied ML systems - balancing latency, observability, and iteration speed • Build scalable services and APIs that bring model outputs to users in trustworthy and intuitive ways • Collaborate across platform, infra, and legal/compliance teams to ensure ML deployments meet standards for safety, fairness, and performance • Establish and evangelize best practices for prompt design, model evaluation, and experimentation across the org

United States
$167.7K - $231.8K / year