Optimization for Reinforcement Learning and Multi-Agent Learning
Foundations and algorithms for policy optimization beyond expected additive rewards, multi-agent learning in structured games, and online learning under limited feedback
My research develops theoretical foundations and algorithms using optimization for learning problems that go beyond the classical paradigm of maximizing expected additive rewards in fixed learning environments. In modern AI systems, objectives may be nonlinear, distributional, multi-sample, or nonconvex with exploitable hidden structure; learning feedback may be limited and revealed online; and environments may contain other agents adapting strategically. These challenges arise in modern AI training and LLM post-training, where models are optimized through imperfect objectives, learned rewards, verifiers, and interactions with users or other agents.
My research is organized around two complementary directions:
Policy Optimization Beyond Expected Additive Rewards: Foundations and optimization algorithms for reinforcement learning and post-training objectives beyond standard expected additive rewards, including general utilities, reward-free RL, risk-sensitive criteria, and LLM post-training.
Multi-Agent Learning in Structured Dynamical Games: Optimization for game dynamics and equilibrium learning in strategic and stateful multi-agent environments including Markov games (stochastic games and beyond, with general utilities) and continuous dynamical games.
For a complete list of my publications, see my Research page below (organized by theme) or Google Scholar (for a list in reverse chronological order).
February 2022–August 2024: Postdoctoral Fellow, ETH Zurich, Department of Computer Science. Worked with Niao He.
Ph.D. in Applied Mathematics and Computer Science, Institut Polytechnique de Paris (Télécom Paris), 2021.
Advised by
Pascal Bianchi and
Walid Hachem.
Engineering Master’s degree in Applied Mathematics and Computer Science, Télécom Paris, 2018.
M.Sc. in Data Science, Université Paris Saclay, 2018.
Here is my CV for more information.
Publications
Keywords: policy gradient methods, general utility RL, convex RL, reward-free RL, LLM post-training.
Keywords: learning in games, game dynamics, multi-agent RL, Markov games with general utilities, Markov potential games, continuous games with state dynamics, game theory.
Keywords: online learning, non-convex optimization, hidden convexity, bandit feedback, regret analysis.
Keywords: Stochastic approximation, Dynamical systems, non-convex stochastic optimization, adaptive gradient methods.
Conferences: NeurIPS, ICML (Gold Reviewer 2026), ICLR, AISTATS, EC.
Journals: Journal of Machine Learning Research (JMLR), Transactions on Machine Learning Research (TMLR), Mathematical Programming, SIAM Journal on Optimization (SIOPT), Journal of Optimization Theory and Applications (JOTA), IEEE Transactions on Automatic Control.
SUTD (2024-2025):
ETH Zurich (2022-2024):
Télécom Paris (2018-2021): Teaching Assistant
Optimization for Machine Learning (graduate level), Statistics (graduate level), Probabilities (undergraduate level).