AI and data programs
10 experts on LLM evaluation in production
These ten engineers and AI leads have worked on evaluating LLM and AI agent systems inside operating companies. A reader can learn how teams build evaluation harnesses, choose quality metrics and decide when an LLM feature is reliable enough to ship.
Professionals to explore
01—10Anil M.
LinkedInExperience: Engineering Manager, AI Agent and LLM Evaluation Infrastructure · Spotify
Engineering Manager, AI Agent and LLM Evaluation Infrastructure at Spotify since August 2025. Other Spotify roles include Engineering Manager, ML AI Infrastructure & Recommendations Delivery (2022–2025) and Engineering Manager, Music Recommendations (2021–2022). The profile says the trickiest part of building AI agents is the harness and the evals, and describes leading two engineering teams at Spotify on that problem.
Dev Jadhav
LinkedInExperience: Tech Lead ML Engineer · ING Nederland
Tech Lead ML Engineer at ING Nederland since November 2024. Other roles on the profile include AI DevOps Architect at ScreenPoint Medical (2023–2024) and Cloud MLOPs Engineer at Lox (2022–2023). The profile headline cites inference reliability and LLM evals and describes building rigorous evaluation harnesses for production LLMs.
Husain Mandsaurwala
LinkedInExperience: Engineering Manager · Shopify
Engineering Manager at Shopify since November 2025; the profile also lists Software Development Manager III at Amazon (2020–2026). The profile headline cites AI infrastructure, model evals and agent platforms, and the summary describes leading AI Infrastructure and Agent Platforms at Shopify on systems that power Sidekick, including owning end-to-end evaluations.
Janita Shah
LinkedInExperience: Software Engineering Manager Amazon Rufus · Amazon
Former Software Engineering Manager Amazon Rufus at Amazon (2025–2026). Other Amazon roles include Engineering Leader Alexa Multi-device Experiences (2020–2025) and Senior Technical Program Manager Alexa (2017–2020). The profile headline lists Gen AI/LLM Evaluation Framework among the areas of work, alongside large scale distributed systems and cloud architecture.
Krishna Pillai
LinkedInExperience: Senior Engineering Manager - Responsible AI , Quality & Infrastructure · Apple
Senior Engineering Manager - Responsible AI, Quality & Infrastructure at Apple since September 2018. The profile headline lists Responsible AI, LLM Evaluations & Quality, and the summary cites expertise in system architecture, cloud platform engineering, developer productivity and ML evaluation.
Maksim Gorev
LinkedInExperience: Senior AI Engineer LLMs, Evals & Text-to-SQL · Triple Whale
Senior AI Engineer LLMs, Evals & Text-to-SQL at Triple Whale (2022–2023), then AI Engineering Manager Applied AI & Agentic Systems there since May 2023. The profile headline cites production agents, evals, context and reliability, and the summary describes building the systems and feedback loops that turn frontier models into dependable production agents while leading AI engineering for Moby, the company's AI agent for commerce.
Prajwal Shreyas
LinkedInExperience: Senior Manager, AI Engineering & Science Lead GenAI · Santander UK
Senior Manager, AI Engineering & Science Lead GenAI at Santander UK since April 2023, after serving there as Principal ML Engineer (2020–2023) and Senior ML Engineer (2018–2020). The profile says this role leads the AI Science and Engineering team, running production AI in a regulated bank where every deployment carries Model Risk sign-off and evals matter as much as the model.
Ravindra Kupatkar
LinkedInExperience: Agentic AI Engineer AI Prompt Engineer · RHB Banking Group
Agentic AI Engineer AI Prompt Engineer at RHB Banking Group since July 2026, after serving as Generative AI Engineer Agentic AI Engineer at Tata Consultancy Services (2023–2026). The profile says the role covers evaluation strategy for the digital division's LLM applications on Azure OpenAI and Azure AI Foundry, including proving an agent behaves before it ships.
Roman Prilepskiy
LinkedInExperience: Team Lead (LLM Evaluation) · Sberbank
Team Lead (LLM Evaluation) at Sberbank since July 2026. Other roles on the profile include LLM Engineer at Avito (2025–2026) and Lead ML developer - MTS AI at MWS AI (2020–2025). The profile summary cites production LLM agents, RAG and search relevance, SFT/LoRA and inference workflows, and prior team leadership of up to 8 specialists.
Tanvi Motwani
LinkedInExperience: Director, Core AI · LinkedIn
Director, Core AI at LinkedIn since February 2026. Other roles on the profile include Director, Applied AI at Block (2024–2026) and Head of Search Engineering at Snap Inc. (2022–2024). The profile cites expertise in search, AI agent evals, LLMs and AI agents, mentions AI Agent Evaluation Science and Platform, and describes leading a machine learning and backend engineering organization that applies AI to products.
Choose the right perspective
Match the person's setting to yours: people who own shared evaluation infrastructure suit platform decisions, while engineers embedded in a single product suit questions about metrics and release checks for one feature. If your launches go through model risk or compliance review, prefer someone whose role sat inside a similar review process.
Questions to take into the conversation
- 01How do you build an evaluation set for an LLM feature, and how do you keep it current as real user traffic changes?
- 02When do you rely on automated scoring such as LLM-as-a-judge, and when do you require human review before a release?
- 03Which production signals tell you an LLM feature is degrading, and what do you do when they fire?
About this directory
This is a professional research starting point based on business profile data retrieved on . Titles and companies reflect that source snapshot and may describe past or present roles. Check the linked profiles for current details. Inclusion does not imply Instant Expert membership or availability.
Request a correction or removal