fal logo
fal

Generative media platform for developers.

Machine Learning Engineer, Reliability

Location

India

Posted

4 days ago

Salary

0

Seniority

Senior

Bachelor Degree3 yrs expEnglishDistributed Systems

Job Description

Machine Learning Engineer, Reliability

fal

• Own availability, latency, and throughput SLOs across a large fleet of generative media model APIs serving production traffic at scale • Build the monitoring, alerting, and observability needed to catch ML-specific failures, output quality degradation, pipeline breakage, model regressions before customers do • Harden model deployment workflows with canary releases, shadow testing, automated rollbacks, and validation gates so new model versions ship safely • Drive the security posture of the model fleet: secure model serving, abuse and misuse detection, rate limiting, and protection against adversarial usage patterns • Operationalize safety systems for generative media, content moderation pipelines, safety classifiers, and guardrails that run reliably at inference time without compromising performance • Lead incident response for model API outages and degradations, run postmortems, and drive the engineering work that prevents recurrence • Improve capacity planning, autoscaling, and GPU fleet efficiency for inference workloads under highly variable traffic • Partner with model and infrastructure teams to make reliability, security, and safety requirements part of how new models get onboarded to the platform

Job Requirements

  • 3+ years of professional experience, with 1 year experience operating production ML or high-scale API systems, ideally with on-call ownership
  • Strong systems fundamentals: distributed systems, networking, observability, and incident management
  • Working knowledge of modern generative models (diffusion, transformers) and their failure modes in production
  • Familiarity with security and safety practices for ML systems ,abuse prevention, content safety, or trust & safety engineering experience is a strong plus
  • A bias toward automation, measurement, and blameless postmortems

Benefits

  • You will have access to our massive GPU cluster for inference and evaluation
  • Some core technologies we use include Python, torch, diffusers, Kubernetes, and the fal Python SDK

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Full TimeRemoteTeam 1,001-5,000

• Design and operate globally distributed backend services • Make architectural decisions on internal service integrations • Manage the full lifecycle: planning, implementation, monitoring • Maintain infrastructure tooling and automation • Contribute to engineering standards and documentation • Evaluate and integrate AI tooling into engineering workflows

Poland
zł28.6K - zł36K / month
Full TimeRemoteTeam 1,001-5,000

• Run the Infrastructure & Services layer for internal data organization • Design, build, maintain, and document internal platforms with focus on automation and reliability • Own stability, availability, and security of services • Engineer systems to be resilient by design • Support services that Product teams depend on • Work directly with Analysts, Scientists, and Engineers • Drive cost optimization and capacity planning • Diagnose hardware faults and dispatch remote hands • Champion automation and engineering best practices • Help shape the structure, tooling, and future direction of the team

Poland
zł23.3K - zł34K / month
Full TimeRemoteTeam 1,001-5,000

• Design, build, and improve monitoring pipelines and observability tooling across globally distributed infrastructure • Define and implement service-level monitoring based on golden signals (latency, traffic, errors, saturation) • Reduce alert fatigue - build meaningful, actionable alerts that engineers trust • Develop and maintain custom exporters, scripts, and integrations for metrics and log collection • Collaborate with the data team on anomaly detection and data-driven operational insights • Understand service signals - know what to measure, why, and what the numbers actually mean

Poland
zł23.3K - zł34K / month
Full TimeRemoteTeam 1,001-5,000

• Ensure content accessibility across a globally distributed edge infrastructure • Design, operate, and improve load balancing and traffic shaping systems at scale • Troubleshoot connection, latency, and accessibility issues end to end • Own HTTP traffic flow - from client to edge to origin and back • Investigate and resolve website and service reachability problems • Work across protocol layers to diagnose issues at the right depth

Poland
zł23.3K - zł34K / month