DeveloperJobs.io
← Back to all jobs

Job Description

Unity Technologies SF is hiring a Staff Machine Learning Engineer to build production-grade, AI-driven game experiences. The role leads research-to-production efforts for computer vision and multi-modal models, spanning cloud, server, and on-device deployments.

Role Overview

In this position, you will help shape the technical roadmap for computer vision and multi-modal AI models, including transformer-based architectures, diffusion networks, vision-language models, and JEPA-style generative approaches. You will translate research prototypes into reliable, well-engineered systems that balance model quality, capability, latency, and cost across target platforms.

Key Responsibilities

  • Set technical vision and roadmap for computer vision and multi-modal AI models, including transformers, diffusion models, vision-language models, and JEPA-style generative architectures.
  • Design and implement models for image and video understanding and generation, including segmentation, detection, and dense prediction, plus multi-modal reasoning over images, text, and 3D inputs.
  • Make architecture, training, data pipeline, and evaluation decisions that balance quality, capability, latency, and cost across cloud, server, and on-device targets.
  • Own the end-to-end path from research to production, including training, fine-tuning, distillation, export, and serving, from cloud GPUs to efficient on-device inference.
  • Partner closely with research scientists to convert novel CV and multi-modal architectures into deployable implementations.
  • Build scalable multi-modal inference systems that handle diverse inputs such as images, video, text, primitives, and metadata, producing outputs ranging from semantic predictions to pixel-level generation.
  • Continuously evaluate and adopt field breakthroughs, including vision-language pretraining and alignment, efficient diffusion methods (for example consistency models and flow matching), efficient attention (such as FlashAttention and linear-attention variants), and tokenization or representation learning for vision.
  • When needed for latency or device constraints, apply compression techniques such as compression, quantization, pruning, and knowledge distillation, and integrate with runtimes including TensorRT, ONNX Runtime, CoreML, and TFLite.
  • Lead and mentor a team of ML engineers by defining engineering best practices, code review standards, and rigorous benchmarking and evaluation methodology.
  • Coordinate with research, platform engineers, product managers, and runtime teams to align ML capabilities with product roadmap needs and platform constraints.
  • Drive a measurement-driven approach by defining KPIs for model quality, accuracy, latency, memory, and cost, and ensuring consistent tracking.

Required Qualifications

  • 6+ years of ML engineering experience with substantial depth in computer vision and/or multi-modal modeling.
  • Production experience with transformer-based and diffusion-based vision models, including examples such as ViT, CLIP/SigLIP-style encoders, Stable Diffusion, and DETR/SAM-style architectures.
  • Comfort across the model lifecycle, including data curation, training and fine-tuning, evaluation, and large-scale serving.
  • Knowledge of efficient attention, diffusion samplers, multi-modal fusion, and vision-language alignment techniques.
  • Strong Python skills and modern deep-learning tooling, including PyTorch, alongside solid software engineering fundamentals.
  • Proven technical leadership experience, including setting direction, influencing cross-functional partners, and mentoring or growing engineers.

Preferred / Additional Experience

  • Experience with world-model, video-generation, or neural rendering pipelines (for example NeRF, 3DGS, or similar).
  • Experience deploying models to constrained or on-device targets, including quantization INT8/INT4/FP16, pruning, distillation, and runtimes such as CoreML, TFLite, and ONNX Runtime.
  • Familiarity with mobile SoC accelerators (for example Apple Neural Engine, Qualcomm Hexagon/Adreno, ARM Mali) or compiler stacks such as MLIR, TVM, or XLA.
  • Contributions to open-source ML frameworks or peer-reviewed CV/ML research publications.
  • Background in real-time graphics or game engine pipelines, including Metal, Vulkan, OpenGL ES.

Technologies

  • Python, PyTorch
  • TensorRT, ONNX Runtime, CoreML, TFLite
  • FlashAttention
  • ViT, CLIP, SigLIP
  • Stable Diffusion, DETR, SAM

Location

Mountain View, CA (onsite)

Salary & Compensation

USD 172,200 - 283,900 per year. Zone details: Zone A: $218,400 - $283,900; Zone B: $194,100 - $252,300; Zone C: $172,200 - $223,900.

Beyond base salary, this role may be eligible for equity awards and participation in company incentive plans, such as annual discretionary bonuses or sales commissions. Final offer amount will depend on geographic location, relevant experience, professional background, and skill set.

Benefits

  • Comprehensive health, life, and disability insurance
  • Commute subsidy
  • Employee stock ownership
  • Competitive retirement/pension plans
  • Generous vacation and personal days
  • Support for new parents through leave and family-care programs
  • Office food snacks
  • Mental Health and Wellbeing programs and support
  • Employee Resource Groups
  • Global Employee Assistance Program
  • Training and development programs
  • Volunteering and donation matching program

Similar Jobs