Michal Valko
Michal is the Founding Researcher at Isara Labs, tenured researcher at Inria, and a lecturer at MVA at ENS Paris-Saclay.
Michal is primarily interested in designing algorithms that would require as little human supervision as possible. He works on methods and settings that are able to deal with minimal feedback, such as deep reinforcement learning, bandit algorithms, self-supervised learning, or self-play. Michal has recently worked on representation learning, world models, and deep (reinforcement) learning algorithms that have some theoretical underpinning. In the past he has also worked on sequential algorithms with structured decisions where exploiting the structure leads to provably faster learning. Michal is now working on a new generation of large language models, in addition to providing algorithmic solutions for their scalable test-time inference, fine-tuning, and alignment.
He received his PhD in 2011 from the University of Pittsburgh, before getting a tenure at Inria in 2012 and co-creating Google DeepMind Paris with Rémi Munos. In 2024, he became a Principal Llama Scientist at Meta (GenAI), building the online reinforcement learning stack and research for Llama 3. In 2025, he joined Isara Labs as a founding researcher.
Talk topics
- Reinforcement learning from human feedback and LLM alignment
- How large language models learn to reason
- Self-supervised representation learning and world models
- Bandit algorithms, exploration, and decision making under uncertainty
Headshots and photos
Click any thumbnail to view full size, or use the Download JPEG button. Photos are free to use in print and online press coverage; please credit the photographer where listed.

















