DeveloperJobs.io
← Back to all jobs

Job Description

Work on Apple’s Human-Centered AI team to evaluate and improve AI systems using data science, model behavior analysis, and qualitative insights.

  • Architect and execute comprehensive evaluation suites for LLMs and multimodal models, surfacing edge cases in multi-step reasoning, factuality, adversarial robustness, safety, and alignment
  • Build deterministic, heuristic, and LLM-assisted evaluation frameworks (including LLM-as-a-judge and reward modeling) to quantify human-perceived quality metrics such as helpfulness and hallucination rates
  • Convert qualitative failure modes into quantifiable loss patterns, programmatic guardrails, and actionable data-mixture adjustments for both training and inference
  • Partner with engineering teams to refine model behavior using evaluation telemetry to guide prompt engineering, Retrieval-Augmented Generation (RAG) strategies, and model fine-tuning
  • Apply advanced ML techniques (embedding-based clustering, representation learning, perturbation analysis) to map error taxonomies and latent failure manifolds
  • Develop robust MLOps workflows to codify evaluation metrics, automate regression testing across model checkpoints, and integrate human-centric assessments into ML CI/CD pipelines
  • Design scalable, distributed inference and processing pipelines (for example, Ray and vLLM) to support high-throughput evaluation, automated annotation, and large-scale output analysis
  • Define quantitative evaluation frameworks capturing nuanced human factors such as trust calibration, conversational state tracking, and interpretability
  • Build automated evaluation pipelines that use LLMs to assess outputs at scale, with optimization for high correlation to human baseline annotations
  • Collaborate across Apple with ML researchers, software developers, and product managers to translate product requirements into reliable and efficient evaluation infrastructure

Requirements

  • Knowledge of human factors, HCI, or cognitive science methodologies as applied to AI system design
  • Bachelor’s or Master’s degree in Computer Science, Machine Learning, Artificial Intelligence, Cognitive Science, or a related technical field
  • 8+ years of relevant industry experience in ML Engineering or Applied Research
  • Advanced proficiency in Python and modern deep learning ecosystems (PyTorch, JAX, Hugging Face)
  • Proven experience building scalable ML inference pipelines, model-evaluation workflows, and structured rating frameworks for large-scale AI systems
  • Strong ability to interpret unstructured model outputs (text, transcripts, embedding spaces) and synthesize qualitative findings into actionable engineering guidance and training objectives
  • Hands-on experience developing, fine-tuning, or evaluating LLMs, multimodal models, and NLP systems
  • Deep familiarity with AI quality metrics, hallucination detection techniques (for example SelfCheckGPT), model alignment (RLHF/DPO), and LLM-as-a-judge frameworks (for example G-Eval and DeepEval)
  • Experience building internal tools or automated pipelines for ML workflows using tools like MLflow and Weights & Biases or similar platforms
  • Strong familiarity with advanced prompt engineering, RAG architectures (vector databases, semantic search), and Fine-Tuning

Technologies

  • Python
  • PyTorch
  • JAX
  • Hugging Face
  • LLM-as-a-judge
  • reward modeling
  • Retrieval-Augmented Generation (RAG)
  • embedding-based clustering
  • representation learning
  • perturbation analysis
  • MLOps
  • CI/CD pipelines
  • Ray
  • vLLM
  • trust calibration
  • SelfCheckGPT
  • RLHF
  • DPO
  • G-Eval
  • DeepEval
  • MLflow
  • Weights & Biases
  • vector databases
  • semantic search
  • Fine-Tuning
  • LLM-assisted evaluation frameworks

Benefits

  • Comprehensive medical and dental coverage
  • Retirement benefits
  • Discounted products and free services
  • Reimbursement for certain educational expenses, including tuition
  • Discretionary restricted stock unit awards
  • Opportunity to purchase Apple stock at a discount via the Employee Stock Purchase Plan
  • Base pay range between $184,700 and $324,800
  • Comprehensive total compensation package may include discretionary bonuses or commission payments as well as relocation

Pay & Benefits

  • Base pay range: $184,700 to $324,800 (depends on skills, qualifications, experience, and location)
  • Employee stock programs: eligibility-dependent discretionary employee stock programs and Employee Stock Purchase Plan participation
  • Benefits include comprehensive medical and dental coverage, retirement benefits, discounted products and free services, and tuition reimbursement for formal education related to career advancement
  • This role might be eligible for discretionary bonuses or commission payments and relocation
  • Note: Apple benefit, compensation, and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program

Similar Jobs