DeveloperJobs.io
← Back to all jobs

Job Description

Apple’s Multimodal Intelligence team is building pipelines, infrastructure, and production systems that translate multimodal foundation models into shipping Apple Intelligence features. This role focuses on end-to-end model delivery, from data curation and training through evaluation and deployment, with an emphasis on meeting on-device and hybrid execution constraints.

Key Responsibilities

  • Design and build pipelines, infrastructure, and production systems that enable multimodal foundation models to ship as Apple Intelligence features.
  • Own end-to-end model delivery, including scaling data curation and training pipelines; fine-tuning and optimizing large multimodal models for on-device and hybrid execution; and implementing reproducible evaluation and regression testing for text and visual understanding.
  • Harden approaches into robust, maintainable production systems while accounting for real constraints such as latency, memory, power, and privacy.
  • Collaborate closely with modeling, platform, hardware, and product engineering teams across Apple to ensure implementation decisions account for future hardware design and product needs.
  • Work broadly with cross-functional partners to deliver strong product outcomes.

Required Qualifications

  • Master’s or PhD, or equivalent practical experience, in Computer Science, Computer Vision, Machine Learning, or a related technical field.
  • Deep expertise in multimodal foundation models with a focus on practical applications.
  • A track record of translating research into practical applications, through published work or industry experience.
  • Applied research experience in at least one major area of model development, including data curation, pre-training, fine-tuning, alignment, or evaluation, particularly for multimodal systems.
  • Experience building large-scale training pipelines, including working with large datasets and scaling models across distributed systems.
  • Experience bridging research ideas with production constraints.
  • Demonstrated deep learning work in at least one area of multimodal systems, such as vision, language, video, or audio.
  • Proficiency in Python and experience with a modern deep learning framework such as PyTorch or JAX.
  • Experience with rapid prototyping, reproduction, and validation of research ideas.
  • Ability to work effectively in a collaborative environment.
  • Ability to communicate analysis results clearly and effectively.
  • BS degree plus a minimum of 3 years of relevant industry experience.

Technical Skills

  • Python
  • PyTorch
  • JAX

Location

Seattle, WA (onsite)

Compensation

The base pay range for this role is USD 142,300 to 263,300 per year. Base pay will depend on skills, qualifications, experience, and location.

Benefits

  • Comprehensive medical and dental coverage
  • Retirement benefits
  • A range of discounted products and free services
  • Reimbursement for certain educational expenses, including tuition
  • Discretionary restricted stock unit awards
  • Participation in Apple’s Employee Stock Purchase Plan allows eligible employees to purchase Apple stock at a discount
  • Opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs
  • Discretionary bonuses or commission payments, as well as relocation (eligibility may apply)

Similar Jobs