Machine Learning Engineer, Human Centered AI - Evaluations & Insights
Job Description
Join Apple’s Human-Centered AI team to build evaluation and insight systems that help foundation models and generative AI deliver reliable, safe, and human-aligned experiences. This role focuses on bridging human perception with model performance through robust evaluation frameworks, scalable MLOps pipelines, and automated analysis that turns findings into actionable guidance for engineering and training.
Location: Seattle, WA (onsite)
Compensation: USD 142,300 - 263,300 per yearly
What you’ll do
- Architect and execute comprehensive evaluation suites for LLMs and multimodal models, surfacing edge cases across multi-step reasoning, factuality, adversarial robustness, safety, and alignment.
- Build deterministic, heuristic, and LLM-assisted evaluation frameworks (including LLM-as-a-judge and reward modeling) to quantify human-perceived quality metrics such as helpfulness and hallucination rates.
- Convert qualitative failure modes into quantifiable loss patterns, programmatic guardrails, and actionable data-mixture adjustments for model training and inference.
- Partner with engineering teams to improve model behavior using evaluation telemetry to guide prompt engineering, Retrieval-Augmented Generation (RAG) strategies, and model fine-tuning.
- Apply advanced ML methods (including embedding-based clustering, representation learning, and perturbation analysis) to map error taxonomies and latent failure manifolds in model outputs.
- Develop robust MLOps workflows that codify evaluation metrics, automate regression testing across model checkpoints, and integrate human-centric assessments into ML CI/CD pipelines.
- Design scalable, distributed inference and processing pipelines (such as Ray and vLLM) to run high-throughput evaluations, automated annotation, and output analysis at scale.
- Define quantitative evaluation frameworks that capture nuanced human factors, including trust calibration, conversational state tracking, and interpretability.
- Build automated evaluation pipelines using LLMs to assess outputs at scale, optimizing for high correlation with human baseline annotations.
- Collaborate across Apple teams, including software engineering, product, research, and responsible AI, to translate product needs into scalable and efficient evaluation infrastructure.
Requirements
- Knowledge of human factors, HCI, or cognitive science methodologies as applied to AI system design.
- 5+ years of relevant industry experience in ML Engineering or Applied Research.
- Advanced proficiency in Python and modern deep learning ecosystems (PyTorch, JAX, Hugging Face).
- Proven experience building scalable ML inference pipelines, model evaluation workflows, and structured rating frameworks for large-scale AI systems.
- Strong ability to interpret unstructured model outputs (text, transcripts, embedding spaces) and synthesize qualitative findings into actionable engineering guidance and training objectives.
- Hands-on experience developing, fine-tuning, or evaluating LLMs, multimodal models, and NLP systems.
- Deep familiarity with AI quality metrics, hallucination detection techniques (including SelfCheckGPT), model alignment approaches (RLHF/DPO), and LLM-as-a-judge frameworks (such as G-Eval and DeepEval).
- Experience building internal tools or automated pipelines for ML workflows using platforms like MLflow and Weights & Biases (or similar).
- Strong familiarity with advanced prompt engineering, RAG architectures (vector databases, semantic search), and fine-tuning.
- Bachelor’s or Master’s degree in Computer Science, Machine Learning, Artificial Intelligence, Cognitive Science, or a related technical field.
Technologies
Python, PyTorch, JAX, Hugging Face, Ray, vLLM, MLflow, Weights & Biases, SelfCheckGPT, RLHF, DPO, G-Eval, DeepEval, LLM-as-a-judge, Retrieval-Augmented Generation (RAG), vector databases, semantic search
Benefits
- Comprehensive medical and dental coverage.
- Retirement benefits.
- Range of discounted products and free services.
- Reimbursement for certain educational expenses (including tuition) for formal education related to advancing your career.
- Opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs.
- Eligible for discretionary restricted stock unit awards.
- Can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan.
- May be eligible for discretionary bonuses or commission payments and relocation.
Preferred qualifications
- Knowledge of human factors, HCI, or cognitive science methodologies as applied to AI system design.