Overview The base model provides capability; post-training turns that capability into a product people want to use. This is a technical, hands-on PM role embedded with applied researchers and ML engineers conducting RLHF, fine-tuning, and preference tuning, translating model outputs into concrete training priorities. You will ensure the end-user voice is reflected in a process driven by metrics and iteration, working with the research team to shape post-training work. Responsibilities - Embed in the post-training loop, working day-to-day with research and applied ML teams during fine-tuning and RLHF cycles, reviewing model outputs and giving structured feedback on quality, tone, and behavior - Define what 'good' means for Copilot models: set priorities for capability, personality, and safety tradeoffs across surfaces (Microsoft 365 Copilot, Copilot Studio agents, Windows) - Run RLHF/preference data campaigns: partner with data and labeling teams to scope what human preference data gets collected and why - Build and maintain behavior evals: translate qualitative judgments into evaluation sets tracked release over release - Own the feedback loop: turn user research, enterprise customer feedback, and production incident learnings into concrete post-training priorities - Make tuning tradeoffs explicit: helpfulness vs. caution, consistency vs. personality, latency/cost vs. quality; drive alignment across research, safety, and product leadership on where the line sits - Represent Responsible AI and enterprise requirements in training
Overview The base model provides capability; post-training turns that capability into a product people want to use. This is a technical, hands-on PM role embedded with applied researchers and ML engineers conducting RLHF, fine-tuning, and preference tuning, translating model outputs into concrete training priorities. You will ensure the end-user voice is reflected in a process driven by metrics and iteration, working with the research team to shape post-training work. Responsibilities - Embed in the post-training loop, working day-to-day with research and applied ML teams during fine-tuning and RLHF cycles, reviewing model outputs and giving structured feedback on quality, tone, and behavior - Define what 'good' means for Copilot models: set priorities for capability, personality, and safety tradeoffs across surfaces (Microsoft 365 Copilot, Copilot Studio agents, Windows) - Run RLHF/preference data campaigns: partner with data and labeling teams to scope what human preference data gets collected and why - Build and maintain behavior evals: translate qualitative judgments into evaluation sets tracked release over release - Own the feedback loop: turn user research, enterprise customer feedback, and production incident learnings into concrete post-training priorities - Make tuning tradeoffs explicit: helpfulness vs. caution, consistency vs. personality, latency/cost vs. quality; drive alignment across research, safety, and product leadership on where the line sits - Represent Responsible AI and enterprise requirements in training
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: Principal Product Manager
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: Principal Product Manager
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
© 2026 JobMatcher. All rights reserved.
© 2026 JobMatcher. All rights reserved.