
large language models, reasoning, fine-tuning, test-time computation, reinforcement learning with human feedback, world models
News RSS feed
News: new Speaking at DTIT 20th Anniversary Event (participation under discussion), October 27, 2026 (tentative), Theatre in Košice, Slovakia.
News: new Espresso grande: Recorded interview in Prague with Romana Gombarčeková.
News: new Speaking at Solar Turbines Digital AI Team: “World Models for Turbines”, October 8, 2026, 10:45-12:15, Solar Turbines EAME Ltd., Za Ženskými domovy 3379, Praha 5-Smíchov, Czechia.
News: new Czech National Bank: Public session on frontier AI capabilities, productivity, labor markets, public institutions, and implications for central banking, with Manoj Pradhan and Iman van Lelyveld; moderated by Stephanie Haffner. Czech National Bank
News: new Speaking at O2 Slovensko 20th Anniversary (participation under discussion), February 4, 2027, evening (tentative), Stará tržnica, Bratislava, Slovakia.
News: new Invited to speak at O2 Slovensko's 20th anniversary celebration in Bratislava on February 4, 2027.
News: new Na jednej vlne: Invited to a one-hour interview in Bratislava; October 15 morning proposed, recording date and time not yet confirmed.
Bio
Michal is the Founding Researcher at Isara Labs, tenured researcher at Inria, and a lecturer at MVA at ENS Paris-Saclay. Michal is primarily interested in designing algorithms that would require as little human supervision as possible. He works on methods and settings that are able to deal with minimal feedback, such as deep reinforcement learning, bandit algorithms, self-supervised learning, or self play. Michal has recently worked on representation learning, world models and deep (reinforcement) learning algorithms that have some theoretical underpinning. In the past he has also worked on sequential algorithms with structured decisions where exploiting the structure leads to provably faster learning. Michal is now working on a new generation of large language models (LLMs), in addition to providing algorithmic solutions for their scalable test-time inference, fine-tuning and alignment. He received his PhD in 2011 from the University of Pittsburgh, before getting a tenure at Inria in 2012 and co-creating Google DeepMind Paris with R. Munos. In 2024, he became a Principal Llama Scientist at Meta, building online reinforcement learning stack and research for Llama 3. In 2025, he joined Isara Labs as a founding researcher.
Selected work
Research threads spanning frontier models, representation learning, bandits, and sparsification.Coming up
Contact
Paris, France
40 avenue Halley
59650 Villeneuve d'Ascq, France
+33 3 59 57 78 01
4, avenue des Sciences
91190 Gif-sur-Yvette, France






