Artificial intelligence
10 AI lab post-training researchers
These ten current and former researchers and engineers have worked on post-training, RLHF and reinforcement learning for language models at AI labs. A reader can learn how post-training researchers approach their work at AI labs from people who have done it.
Professionals to explore
01—10Aliaksei Severyn
LinkedInExperience: Research Director | Principal Research Scientist · Google DeepMind
Research Director and Principal Research Scientist at Google DeepMind since 2024; the profile describes him as Gemini post-training and user feedback lead, including leading post-training work for the Gemini app. Earlier research scientist roles at Google from 2016 to 2024, most recently Senior Staff Research Scientist.
Devamanyu Hazarika
LinkedInExperience: Research Scientist · Meta
Research Scientist at Meta since 2025; the profile says he works on post-training in the Llama research team within Meta Superintelligence Labs. Former Senior Applied Scientist at Amazon (2024-2025), where the profile describes work on model alignment for the Amazon Nova models.
Haochen Zhang
LinkedInExperience: Member of Technical Staff · Anthropic
Member of Technical Staff at Anthropic since 2024; the profile says he works on reinforcement learning and post-training of Claude models. He also lists PhD in Computer Science at Stanford University (2019-2024), where the profile describes work on the theoretical foundations of deep learning.
Haoran Li
LinkedInExperience: Research TL @ GenAI · Meta
Former research tech lead in GenAI at Meta (2017-2024), where the profile says he led LLM post-training for character chatbots and AI Studio, covering fine-tuning, RLHF, evaluation and the user data flywheel. Head of Post-training and Data (Research) at Character.ai since 2025.
Jiahui Li
LinkedInExperience: Senior Research Engineer, AGI Post-Training · Amazon
Senior Research Engineer, AGI Post-Training at Amazon since 2025, after serving as Research Engineer, AGI Post-Training (2023-2025). The profile describes building large-scale asynchronous and off-policy RL systems for foundation model post-training. Earlier a Software Engineer in Alexa AI at Amazon (2020-2023).
Jyh-Jing Hwang
LinkedInExperience: Research Scientist · Google DeepMind
Research Scientist at Google DeepMind since 2025, where the profile says he focuses on Gemini post-training research and agentic reinforcement learning. Former Research Scientist and tech lead manager at Waymo (2020-2025), and earlier a Graduate Student Researcher at the International Computer Science Institute (2016-2020).
Rohan Kshirsagar
LinkedInExperience: Member of Technical Staff · OpenAI
Member of Technical Staff at OpenAI since 2024; the profile headline places him in post-training research, with work on evals, hack robustness, RL interpretability and user signal reward modeling. Former Staff Machine Learning Engineer at Airbnb (2018-2022) and Senior Software Engineer in the NLP Group at Bloomberg (2013-2016).
Shunyu Yao
LinkedInExperience: Senior Staff Research Scientist · Google DeepMind
Senior Staff Research Scientist at Google DeepMind since 2025, working on Gemini post-training research according to the profile. Former Member of Technical Staff at Anthropic (2024-2025), where the profile says he was a research scientist on reinforcement learning. PhD work at Stanford University (2019-2024).
Yannis Flet-Berliac
LinkedInExperience: Research Scientist · Cohere
Research Scientist at Cohere since 2024, where the profile says he focuses on large language model post-training and reinforcement learning research. Former Research Scientist at InstaDeep (2023-2024) and postdoctoral scholar at Stanford University (2022-2023) working on offline and deep reinforcement learning.
Yunzhe Tao
LinkedInExperience: Member of Technical Staff · Microsoft AI
Member of Technical Staff at Microsoft AI since 2026; the profile headline places him in post-training on the superintelligence team. Former Founding Research Scientist at Anuttacon (2024-2026), where he led post-training of model behavior, and Staff Research Scientist at ByteDance (2021-2024).
Choose the right perspective
Match the person's role type to your question, since research scientists, research engineers building RL systems and research leads each see post-training differently. Consider which lab and which years are closest to your decision, since several people have worked at more than one lab, and whether a current or former employee suits it.
Questions to take into the conversation
- 01How do post-training teams decide which behaviors to target with reinforcement learning?
- 02What makes a reward signal from user feedback useful or misleading?
- 03How do research engineers build RL systems that scale for large models?
Reaching AI lab post-training researchers
How can I contact one of these AI lab post-training researchers?
Pick one of the 10 people on this page and choose "Book a paid call", or describe your project to find others. You offer a fee for a 15-minute call or a written answer, Instant Expert finds the person's work email and sends the invitation, and they decide whether to accept. The list covers 7 companies, including Google DeepMind, Meta and Anthropic, based on profile data retrieved on October 9, 2026. Being listed here doesn't mean someone has agreed to take calls.
How much does it cost to reach AI lab post-training researchers?
You choose the offer, starting at $5 per person. It includes Instant Expert's 20% fee, so an offer that pays the person $100 costs you $125. You're charged only when the person books the call or sends the answer.
What if they don't reply?
You pay nothing. Instant Expert sends follow-up reminders, and if the person hasn't booked or answered within 7 days, the request expires and any hold on your card is released. You can invite several of the 10 people on this list at once and cap your total spend, so you pay only for the ones who accept.
Can an AI agent ask one of these AI lab post-training researchers a question?
Yes. An agent with a USDC wallet on Base can ask one question without an Instant Expert account. It names the person, for example by the LinkedIn URL on this page, and pays per ask over x402 (HTTP 402). The person answers in writing or by voice note, and if nobody answers within 7 days, the payment goes back to the wallet automatically. x402 docs
About this directory
This is a professional research starting point based on business profile data retrieved on . Titles and companies reflect that source snapshot and may describe past or present roles. Check the linked profiles for current details. Inclusion does not imply Instant Expert membership or availability.
Request a correction or removal