DeveloperJobs.io
← Back to all jobs

Job Description

This hybrid Machine Learning Engineering and Site Reliability Engineering role focuses on reliability, security, and safety for fal’s generative media model API fleet.

Responsibilities

  • Own availability, latency, and throughput SLOs for a large fleet of generative media model APIs serving production traffic at scale
  • Design and run monitoring, alerting, and observability to detect ML-specific issues, including output quality degradation, pipeline breakage, and model regressions
  • Harden model deployment workflows using canary releases, shadow testing, automated rollbacks, and validation gates for safer model version shipping
  • Improve the security posture of the model fleet with secure model serving, abuse and misuse detection, rate limiting, and protection against adversarial usage patterns
  • Operationalize safety systems for generative media, including content moderation pipelines, safety classifiers, and guardrails that run reliably at inference time without sacrificing performance
  • Lead incident response for model API outages and degradations, perform postmortems, and drive engineering changes to prevent recurrence
  • Strengthen capacity planning, autoscaling, and GPU fleet efficiency for inference workloads with highly variable traffic
  • Partner with model and infrastructure teams so reliability, security, and safety requirements are built into new model onboarding to the platform

Requirements

  • 5+ years of professional experience
  • 2+ years operating production ML or high-scale API systems, ideally with on-call ownership
  • Experience working with and supporting diffusion models in production
  • Strong systems fundamentals across distributed systems, networking, observability, and incident management
  • Working knowledge of modern generative models (diffusion, transformers) and their production failure modes
  • Familiarity with security and safety practices for ML systems (abuse prevention, content safety, or trust and safety experience is a plus)
  • Automation-oriented mindset with emphasis on measurement and blameless postmortems

Tech & Tools

  • Python
  • torch
  • diffusers
  • Kubernetes
  • fal Python SDK

Location

  • Remote - APAC
  • Role will need to be based in India, Australia, or New Zealand

Employment

  • Full time
  • Department: EngineeringML

Working Environment

  • Access to a massive GPU cluster for inference and evaluation
  • Work alongside a team focused on quickly iterating on and deploying new AI breakthroughs while maintaining reliability

Similar Jobs