
fal
Remote Jobs
Generative media platform for developers.
13 Jobs
• Resolve technical issues and provide advanced support directly to customers, including support for fal’s platform (APIs, UI issues, and troubleshooting errors). • Support users across multiple products via email, chat, and Slack. • Troubleshoot integration issues, including authentication problems (OAuth, API keys), HTTP errors, malformed requests, rate limits, and API misconfigurations. • Analyze API logs, error messages, and request/response payloads to identify root causes. • Manage support tickets by responding within SLA timeframes, escalating complex issues appropriately, and maintaining detailed case records. • Reproduce, escalate, and document bugs or edge cases in collaboration with engineering. • Provide structured feedback to engineering teams regarding platform reliability, performance bottlenecks, and customer-reported issues, serving as an internal advocate for customer pain points and product improvement. • Assist with testing and validation of new features, releases, and infrastructure changes before production deployment. • Write and maintain technical content, including use case guides, how-to examples, FAQs, solutions for common errors, and documentation of issues and resolutions for the knowledge base. • Improve developer documentation to make integration as self-serve as possible.
• Build and maintain Python fleet tracking system that manages the full lifecycle of servers including contracting and procurement, target use, pricing, availability, health, RMAs, etc • Build server management tooling that automates provisioning, health checks, GPU diagnostics, recovery and alerting • Create and maintain metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk failures, network issues, thermals) • Leverage AI to an extreme level to build tools and automate alerting and recovery • Implement and enforce OS-level security: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation • Manage and optimize distributed and local storage systems supporting model weights, checkpoints, and ephemeral scratch: NVMe arrays, NFS, parallel file systems, and object storage • Tune Linux systems for AI workloads: kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stack optimization (NVIDIA drivers, CUDA, container runtimes) • Develop a suite of automated error detection and recovery processes • Work with partners to solve technical issues
• Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads • Build and maintain CI/CD pipelines and deployment infrastructure • Leverage AI to an extreme level to automate analysis and resolution of production issues, and improve software development speed, reliability and maintainability • Build dashboards, alerting, and anomaly detection across our systems • Define and enforce SLOs and build out incident response processes • Manage and improve our networking, load balancing, and service mesh configurations • Drive reliability improvements across the stack through automation, runbooks, and chaos engineering
• Own availability, latency, and throughput SLOs across a large fleet of generative media model APIs serving production traffic at scale • Build the monitoring, alerting, and observability needed to catch ML-specific failures, output quality degradation, pipeline breakage, model regressions before customers do • Harden model deployment workflows with canary releases, shadow testing, automated rollbacks, and validation gates so new model versions ship safely • Drive the security posture of the model fleet: secure model serving, abuse and misuse detection, rate limiting, and protection against adversarial usage patterns • Operationalize safety systems for generative media, content moderation pipelines, safety classifiers, and guardrails that run reliably at inference time without compromising performance • Lead incident response for model API outages and degradations, run postmortems, and drive the engineering work that prevents recurrence • Improve capacity planning, autoscaling, and GPU fleet efficiency for inference workloads under highly variable traffic • Partner with model and infrastructure teams to make reliability, security, and safety requirements part of how new models get onboarded to the platform
• Build our core Python/Rust platform: request routing, AI workload orchestration, scheduling, GPU autoscaling, large scale file storage, queueing, etc • Produce forward designs for platform evolution as we scale to 100x current traffic and need to provide low latency across the world • Leverage AI to an extreme level to automate the mundane parts of building complex but reliable systems • Profile and tune low level CPU and memory performance
• Monitor health and performance of InfiniBand and Ethernet fabrics: switches, HCAs, transceivers, links. • Investigate and resolve fabric issues: connectivity, congestion, performance regressions. • Support fabric bring-up alongside DC ops and customer-facing teams. • Run maintenance and upgrades on switches and control plane components. • Partner with cluster ops on cross-domain incidents where the line between compute and network is blurry. • Improve the tooling and runbooks so the next incident resolves faster than the last.
• Build and lead the Fleet Reliability team: hire, develop, retain • Own 24/7 coverage for node provisioning, validation, and triage • Drive the automation roadmap: event-driven remediation, self-healing, observability • Define and enforce the SLAs that keep production GPUs serving traffic • Set the culture: how the team keeps score, how they communicate, how they grow
• Provision, validate, and triage GPU nodes across B300, H200, and H100 clusters • Troubleshoot hardware and software issues across compute, network, and storage • Monitor fleet health, take remediation action, push fixes upstream when needed • Write the runbooks. Improve the ones that exist. Delete the ones that don't work
• Design, fit-out, and commission white space across owned and colo sites: PDUs, CDUs, busway, in-row cooling, high-density power. • Build the smart hands and break-fix teams. • Run a multi-site portfolio across stages: build, commission, ramp, live • Hold vendors and integrators to schedule and spec: Supermicro, NVIDIA, cooling vendors, cable plant subs, the GC • Own the technical calls on white space: rack layout, cooling topology, cable plant, power redundancy, serviceability • Partner with network eng on InfiniBand/Interconnects • Establish standards that scale, so site #10 doesn't need you on every commissioning call.
Role Description As a Senior Data Scientist for Go-to-Market at fal, you will be the analytical backbone of our revenue organization. Embedded directly with the GTM function, you will own the metrics that tell us how our pipeline is performing, how our sales team is executing, and where the next dollar of revenue is most likely to come from. This is a high-leverage, high-visibility role. GTM leadership will look to you for the answers - on rep performance, quota attainment, pipeline health, AM coverage, and segment economics - and you'll shape both the questions we ask and the systems we build to answer them. You'll partner closely with Product Intelligence and Data Engineering as part of fal's center-of-excellence data team, while acting as the embedded specialist for everything GTM. Key responsibilities: - Own the metrics and reporting that drive sales execution at fal: pipeline health, quota attainment, conversion rates, sales cycle, AM coverage, and segment-level economics. - Partner directly with the sales leadership and rev-ops to translate strategic questions into measurement frameworks and weekly operating cadences. - Build the data foundations that make GTM analytics first-class: account-manager tracking, territory and quota models, opportunity-stage instrumentation, and rep-level performance views. - Run rigorous deep-dives - win/loss, lead source effectiveness, segmentation, expansion drivers - that change how GTM allocates time and headcount. - Shape the CRM, internal tooling, and downstream data models that everyone in GTM depends on, and set the standards for how GTM data is captured, joined, and trusted. Qualifications - 5+ years of experience in data science or analytics roles, with at least 3+ years specifically in GTM, RevOps, or sales analytics at an enterprise or B2B SaaS company. - Advanced SQL and proficiency in Python for analytics, modeling, and experimentation. - Deep familiarity with Salesforce data, pipeline mechanics, quota and territory models, and the operational rhythms of an enterprise sales org. - Strong proficiency with dbt and modern analytics stacks; comfort partnering with data engineering on production-grade pipelines. - Proven ability to operate as an embedded analytics partner - building trust with sales leadership and translating ambiguous strategic asks into clear measurement. - Demonstrated bias for action; you set up the dashboard, write the SQL, and ship the framework yourself when needed. - A track record of shipping data products that change behavior, not just dashboards that get viewed once. Requirements - Nice to have: - Experience supporting both self-serve/PLG and enterprise sales motions. - Experience working on developer-facing or API products. - Familiarity with usage-based pricing and consumption-driven sales motions. - Early-stage or fast-scaling environment experience. Benefits - Compensation: $180,000-225,000 plus equity + benefits (This range is across 2 levels Senior and Staff). - Location: San Francisco, CA (willing to consider remote for Senior and Staff levels). - Interesting and challenging work. - A lot of learning and growth opportunities. - We are currently hiring in downtown San Francisco. - We offer relocation assistance to San Francisco. - Health, dental, and vision insurance (US). - Regular team events and offsites.
3more opportunities are still waiting for you.Log in now and take your next shot before someone else does.