Senior Applied Scientist with deep expertise in Applied AI, Machine Learning, and Deep Learning to deliver innovative, user-facing AI experiences at scale. This role shapes the future of enterprise productivity through Large Language Models (LLMs) and AI-powered experiences in Microsoft 365 Copilot. The role is based in London, United Kingdom with a hybrid work arrangement.
Senior Applied Scientist with deep expertise in Applied AI, Machine Learning, and Deep Learning to deliver innovative, user-facing AI experiences at scale. This role shapes the future of enterprise productivity through Large Language Models (LLMs) and AI-powered experiences in Microsoft 365 Copilot. The role is based in London, United Kingdom with a hybrid work arrangement.
Develop scientific approaches to evaluate and improve LLM reasoning, grounding, retrieval, and agentic behaviors using offline benchmarks, online experimentation, and large-scale telemetry analysis.
Drive complex technical investigations, model evaluations, and experimentation to improve quality, reliability, latency, cost efficiency, and user experience.
Define and execute rigorous evaluation methodologies, offline benchmarks, online experimentation frameworks, and data-driven approaches for measuring model and product performance.
Partner with engineering and product teams to translate research into scalable, production-ready systems used by millions of customers.
Deeply analyze model behavior, failure modes, grounding effectiveness, reasoning quality, tool use, retrieval performance, and orchestration outcomes to identify opportunities for improvement.
Develop approaches for agentic workflows, multi-step reasoning, planning, retrieval-augmented generation (RAG), tool orchestration, and reinforcement-learning-based optimization.
Review technical designs, experimental results, and implementation approaches to ensure high standards of scientific rigor, reproducibility, an
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: Senior Applied Scientist
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
Develop scientific approaches to evaluate and improve LLM reasoning, grounding, retrieval, and agentic behaviors using offline benchmarks, online experimentation, and large-scale telemetry analysis.
Drive complex technical investigations, model evaluations, and experimentation to improve quality, reliability, latency, cost efficiency, and user experience.
Define and execute rigorous evaluation methodologies, offline benchmarks, online experimentation frameworks, and data-driven approaches for measuring model and product performance.
Partner with engineering and product teams to translate research into scalable, production-ready systems used by millions of customers.
Deeply analyze model behavior, failure modes, grounding effectiveness, reasoning quality, tool use, retrieval performance, and orchestration outcomes to identify opportunities for improvement.
Develop approaches for agentic workflows, multi-step reasoning, planning, retrieval-augmented generation (RAG), tool orchestration, and reinforcement-learning-based optimization.
Review technical designs, experimental results, and implementation approaches to ensure high standards of scientific rigor, reproducibility, an
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: Senior Applied Scientist
See how this role matches your resume
Sign in to get an AI match score and personalized feed.