Apple is seeking a Machine Learning Evaluation Engineer to design and scale evaluation capabilities for our next generation of AI-powered experiences. You will design automated evaluation frameworks and pipelines for LLMs, Generative AI, Conversational AI, and Agentic AI products. Responsibilities include defining evaluation methodologies and quality metrics across dimensions such as accuracy, relevance, groundedness, completeness, consistency, instruction following, and task completion; building and maintaining high-quality evaluation datasets, including golden datasets, benchmark sets, regression suites, adversarial scenarios, and production-derived test sets; developing Auto Eval capabilities to rapidly evaluate models, prompts, retrieval systems, agents, and end-to-end AI experiences; designing and implementing model-based evaluation approaches (including LLM-as-a-Judge) with calibration and validation methodologies; developing Human-in-the-Loop evaluation approaches for complex or subjective quality dimensions where automated evaluation alone is insufficient; defining evaluation rubrics, annotation guidelines, grading criteria, and quality standards in partnership with product teams, domain experts, and annotation teams; building mechanisms to calibrate automated evaluators against human judgment and measure evaluator consistency and reliability; evaluating end-to-end AI systems, including retrieval, context construction, prompts, model responses, tool use, APIs, and downstream product experiences; developing evaluation methodologies for multi-turn conversations, perso
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: Machine Learning Evaluation Engineer
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
Apple is seeking a Machine Learning Evaluation Engineer to design and scale evaluation capabilities for our next generation of AI-powered experiences. You will design automated evaluation frameworks and pipelines for LLMs, Generative AI, Conversational AI, and Agentic AI products. Responsibilities include defining evaluation methodologies and quality metrics across dimensions such as accuracy, relevance, groundedness, completeness, consistency, instruction following, and task completion; building and maintaining high-quality evaluation datasets, including golden datasets, benchmark sets, regression suites, adversarial scenarios, and production-derived test sets; developing Auto Eval capabilities to rapidly evaluate models, prompts, retrieval systems, agents, and end-to-end AI experiences; designing and implementing model-based evaluation approaches (including LLM-as-a-Judge) with calibration and validation methodologies; developing Human-in-the-Loop evaluation approaches for complex or subjective quality dimensions where automated evaluation alone is insufficient; defining evaluation rubrics, annotation guidelines, grading criteria, and quality standards in partnership with product teams, domain experts, and annotation teams; building mechanisms to calibrate automated evaluators against human judgment and measure evaluator consistency and reliability; evaluating end-to-end AI systems, including retrieval, context construction, prompts, model responses, tool use, APIs, and downstream product experiences; developing evaluation methodologies for multi-turn conversations, perso
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: Machine Learning Evaluation Engineer
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
5 active roles from this employer in the JobMatcher catalog.
Data Engineer
Evans Transportation
Parts Sales Manager
AutoZone
Retail Team Lead (Part-Time)
Office Depot
Airline Program Manager
Letronics
Commercial Sales Manager
AutoZone
Beauty Advisor (Inside Sales)
Sally Beauty
Parts Sales Manager
AutoZone
Quality Engineer
LotusWorks
Business Intelligence Analyst
Christian Brothers Automotive
Beauty Advisor
Sally Beauty
5 active roles from this employer in the JobMatcher catalog.
Data Engineer
Evans Transportation
Parts Sales Manager
AutoZone
Retail Team Lead (Part-Time)
Office Depot
Airline Program Manager
Letronics
Commercial Sales Manager
AutoZone
Beauty Advisor (Inside Sales)
Sally Beauty
Parts Sales Manager
AutoZone
Quality Engineer
LotusWorks
Business Intelligence Analyst
Christian Brothers Automotive
Beauty Advisor
Sally Beauty
© 2026 JobMatcher. All rights reserved.
© 2026 JobMatcher. All rights reserved.