DeveloperJobs.io
← Back to all jobs

Job Description

Apple is building next-generation intelligent search and AI experiences, with systems designed to understand user intent and context while preserving privacy. In this onsite role in Santa Clara, you will help take transformer-based language models and retrieval systems from research into production, with a strong focus on efficient deployment on-device.

This Senior Machine Learning Engineer position supports privacy-conscious AI by combining semantic retrieval, ranking, and retrieval-augmented generation with model training, fine-tuning, optimization, and deployment. The work spans evaluation methodology, scalable experimentation pipelines, and techniques to transfer capabilities from large foundation models into compact on-device models.

What you’ll do

  • Build semantic retrieval, embedding, reranking, and retrieval-augmented generation systems, plus models for query understanding, intent prediction, personalization, retrieval, and ranking.
  • Analyze search relevance and user behavior to develop evaluation methodologies, offline benchmarks, and online metrics for retrieval quality, ranking, personalization, and language model performance.
  • Create scalable experimentation and evaluation pipelines for LLMs and search models, including measures for model quality, robustness, latency, efficiency, and end-to-end product metrics.
  • Design, train, fine-tune, distill, and optimize transformer-based language models and foundation models for efficient on-device deployment.
  • Develop LLM fine-tuning and post-training approaches such as supervised fine-tuning, instruction tuning, preference optimization, parameter-efficient fine-tuning, and task-specific adaptation.
  • Research and prototype on-device generative AI approaches including knowledge distillation, model compression, quantization, pruning, and low-latency inference.
  • Transfer capabilities from large foundation models into compact on-device models while balancing quality, latency, memory footprint, power consumption, and compute constraints.
  • Partner with engineers, researchers, product managers, and designers to move AI capabilities from research into production, including work spanning foundation models, multimodal AI, agentic retrieval, and personalized intelligence.

Minimum qualifications

  • Master degree in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
  • 5+ years of industry or research experience developing machine learning systems.
  • Background in machine learning, deep learning, natural language processing, information retrieval, search, recommender systems, or generative AI.
  • Experience training, fine-tuning, or deploying transformer-based models and large language models.
  • Experience with modern deep learning architectures and techniques including transformers, embeddings, representation learning, and neural ranking.
  • Programming skills in Python and/or C/C++, with experience building production-quality software using modern machine learning frameworks such as PyTorch, JAX, or TensorFlow.
  • Ability to work onsite in Cupertino, California, in accordance with Apple's applicable work policies.

Additional requirements

  • Master's or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
  • Experience optimizing machine learning models for resource-constrained environments, including knowledge distillation, model compression, quantization, and pruning.
  • Experience with on-device machine learning or edge AI, or mobile inference frameworks, optimizing for latency, memory, compute, and power constraints.
  • Experience distilling capabilities from large foundation models into small language models or task-specific models for efficient inference.
  • Experience building retrieval-augmented generation, vector search, embedding retrieval, neural reranking, or semantic search systems.
  • Experience with query understanding, query rewriting, intent classification, personalized retrieval, learning-to-rank, or recommendation models.
  • Experience with transformer architectures and foundation model families such as BERT, T5, Llama, Gemma, Mistral, or related architectures.
  • Experience evaluating language models, designing AI quality metrics, and building automated and human-in-the-loop evaluation pipelines.
  • Experience building large-scale production search, recommendation, personalization, or generative AI systems.
  • Familiarity with multimodal foundation models, tool use, agentic AI, or agentic retrieval systems.
  • Strong understanding of tradeoffs among model quality, latency, memory, power consumption, privacy, and reliability for production on-device AI systems.
  • Ability to prototype new ideas, conduct rigorous experiments, solve ambiguous technical problems, and translate research advances into production-quality machine learning solutions.

Technologies

  • Python, C/C++, PyTorch, JAX, TensorFlow, transformers, BERT, T5, Llama, Gemma, Mistral

Pay & benefits

  • Base pay range: USD 184,700 - 324,800 per year, depending on skills, qualifications, experience, and location.
  • Opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs.
  • Eligible for discretionary restricted stock unit awards and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan.
  • Comprehensive medical and dental coverage.
  • Retirement benefits.
  • Range of discounted products and free services.
  • Reimbursement for certain educational expenses, including tuition, for formal education related to advancing your career at Apple.
  • May be eligible for discretionary bonuses or commission payments as well as relocation.
  • Note: Benefit, compensation, and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Similar Jobs