Senior Machine Learning Engineer, Proactive
Job Description
Apple is building next-generation intelligent search and AI experiences, with systems designed to understand user intent and context while preserving privacy. In this onsite role in Santa Clara, you will help take transformer-based language models and retrieval systems from research into production, with a strong focus on efficient deployment on-device.
This Senior Machine Learning Engineer position supports privacy-conscious AI by combining semantic retrieval, ranking, and retrieval-augmented generation with model training, fine-tuning, optimization, and deployment. The work spans evaluation methodology, scalable experimentation pipelines, and techniques to transfer capabilities from large foundation models into compact on-device models.
What you’ll do
- Build semantic retrieval, embedding, reranking, and retrieval-augmented generation systems, plus models for query understanding, intent prediction, personalization, retrieval, and ranking.
- Analyze search relevance and user behavior to develop evaluation methodologies, offline benchmarks, and online metrics for retrieval quality, ranking, personalization, and language model performance.
- Create scalable experimentation and evaluation pipelines for LLMs and search models, including measures for model quality, robustness, latency, efficiency, and end-to-end product metrics.
- Design, train, fine-tune, distill, and optimize transformer-based language models and foundation models for efficient on-device deployment.
- Develop LLM fine-tuning and post-training approaches such as supervised fine-tuning, instruction tuning, preference optimization, parameter-efficient fine-tuning, and task-specific adaptation.
- Research and prototype on-device generative AI approaches including knowledge distillation, model compression, quantization, pruning, and low-latency inference.
- Transfer capabilities from large foundation models into compact on-device models while balancing quality, latency, memory footprint, power consumption, and compute constraints.
- Partner with engineers, researchers, product managers, and designers to move AI capabilities from research into production, including work spanning foundation models, multimodal AI, agentic retrieval, and personalized intelligence.
Minimum qualifications
- Master degree in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
- 5+ years of industry or research experience developing machine learning systems.
- Background in machine learning, deep learning, natural language processing, information retrieval, search, recommender systems, or generative AI.
- Experience training, fine-tuning, or deploying transformer-based models and large language models.
- Experience with modern deep learning architectures and techniques including transformers, embeddings, representation learning, and neural ranking.
- Programming skills in Python and/or C/C++, with experience building production-quality software using modern machine learning frameworks such as PyTorch, JAX, or TensorFlow.
- Ability to work onsite in Cupertino, California, in accordance with Apple's applicable work policies.
Additional requirements
- Master's or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
- Experience optimizing machine learning models for resource-constrained environments, including knowledge distillation, model compression, quantization, and pruning.
- Experience with on-device machine learning or edge AI, or mobile inference frameworks, optimizing for latency, memory, compute, and power constraints.
- Experience distilling capabilities from large foundation models into small language models or task-specific models for efficient inference.
- Experience building retrieval-augmented generation, vector search, embedding retrieval, neural reranking, or semantic search systems.
- Experience with query understanding, query rewriting, intent classification, personalized retrieval, learning-to-rank, or recommendation models.
- Experience with transformer architectures and foundation model families such as BERT, T5, Llama, Gemma, Mistral, or related architectures.
- Experience evaluating language models, designing AI quality metrics, and building automated and human-in-the-loop evaluation pipelines.
- Experience building large-scale production search, recommendation, personalization, or generative AI systems.
- Familiarity with multimodal foundation models, tool use, agentic AI, or agentic retrieval systems.
- Strong understanding of tradeoffs among model quality, latency, memory, power consumption, privacy, and reliability for production on-device AI systems.
- Ability to prototype new ideas, conduct rigorous experiments, solve ambiguous technical problems, and translate research advances into production-quality machine learning solutions.
Technologies
- Python, C/C++, PyTorch, JAX, TensorFlow, transformers, BERT, T5, Llama, Gemma, Mistral
Pay & benefits
- Base pay range: USD 184,700 - 324,800 per year, depending on skills, qualifications, experience, and location.
- Opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs.
- Eligible for discretionary restricted stock unit awards and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan.
- Comprehensive medical and dental coverage.
- Retirement benefits.
- Range of discounted products and free services.
- Reimbursement for certain educational expenses, including tuition, for formal education related to advancing your career at Apple.
- May be eligible for discretionary bonuses or commission payments as well as relocation.
- Note: Benefit, compensation, and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.