
large language models, reasoning, fine-tuning, test-time computation, reinforcement learning with human feedback, world models
News RSS feed
News: new Technical report: DiG-bench: Discovery in Games, paper a benchmark for scientific discovery through interaction and experimentation!
News: new ASAI Stream features my AI Personality 2026 recognition for Social Impact of the Year: weekly Slovak AI ecosystem overview!
News: Invited talk at the Armenia LLM Summer School in Yerevan, Armenia (Aug 3-7)!
News: Interview in Grand Magazine: Majme v umelej inteligencii nášho partnera, nie roz...!
News: Speaking at Future Week 2026 in Bergen, Norway!
News: Talk at Krafton AI in Seoul, South Korea!
News: Talk at the Global AI Show: AI 2030 in Riyadh, Saudi Arabia (June 29-30)!
Bio
Michal is the Founding Researcher at Isara Labs, tenured researcher at Inria, and a lecturer at MVA at ENS Paris-Saclay. Michal is primarily interested in designing algorithms that would require as little human supervision as possible. He works on methods and settings that are able to deal with minimal feedback, such as deep reinforcement learning, bandit algorithms, self-supervised learning, or self play. Michal has recently worked on representation learning, world models and deep (reinforcement) learning algorithms that have some theoretical underpinning. In the past he has also worked on sequential algorithms with structured decisions where exploiting the structure leads to provably faster learning. Michal is now working on a new generation of large language models (LLMs), in addition to providing algorithmic solutions for their scalable test-time inference, fine-tuning and alignment. He received his PhD in 2011 from the University of Pittsburgh, before getting a tenure at Inria in 2012 and co-creating Google DeepMind Paris with R. Munos. In 2024, he became a Principal Llama Scientist at Meta, building online reinforcement learning stack and research for Llama 3. In 2025, he joined Isara Labs as a founding researcher.
Selected work
Research threads spanning frontier models, representation learning, bandits, and sparsification.
BYOL & bootstrapped learning
Bootstrapped self-supervision that grew from images into video, graphs, neural activity, and exploration.
BYOL · BraVe · BGRL · SwapVAE · BYOL-Explore
Nash-MD / NLHF
Preference learning as a game, connecting alignment to Nash-equilibrium algorithms.
NLHF · Nash-MD · preference learning
IX / implicit exploration
Implicit exploration replaces explicit exploration mixing with a biased loss estimator, giving sharp guarantees under partial and side-observation feedback.
side observations · data-dependent bounds · contextual bandits
Bandits & structured exploration
A long line on learning faster from structure: graphs, side observations, combinatorial actions, and Bayesian exploration.
Spectral · graph bandits · best-arm ID · semi-bandits · Thompson sampling
SQUEAK & sparsification
Single-pass sparsification of kernel matrices and graphs using ridge leverage scores, with scalable distributed variants.
Kernel approximation · ridge leverage scores · spectral sparsification
Coming up
Contact
Paris, France
40 avenue Halley
59650 Villeneuve d'Ascq, France
+33 3 59 57 78 01
4, avenue des Sciences
91190 Gif-sur-Yvette, France




































