My research lies at the intersection of optimization, reinforcement learning, and learning in games. I develop optimization foundations and algorithms for sequential decision-making, with a focus on richer objectives, strategic interactions, and limited feedback, motivated by emerging challenges in modern AI systems.
I am currently a Research Fellow at Singapore University of Technology and Design, working with Georgios Piliouras and Antonios Varvitsiotis. Previously, I was a postdoctoral fellow at ETH Zurich working with Niao He. I received my PhD in mathematics and computer science from Institut Polytechnique de Paris (Télécom Paris), advised by Pascal Bianchi and Walid Hachem.
Optimization for Sequential Decision-Making
My work studies optimization and learning for sequential decision-making when objectives go beyond expected cumulative rewards, agents interact strategically, or feedback is limited. My research is organized around three complementary directions:
Policy Optimization Beyond Expected Additive Rewards: Foundations and optimization algorithms for reinforcement learning with objectives beyond standard expected additive rewards, including general utilities and LLM post-training.
Learning and Optimization in Structured Games: Foundations and algorithms for strategic and stateful multi-agent environments, with a focus on equilibria and learning dynamics.
Online Learning with Structured Objectives under Limited Feedback: Learning and optimization by exploiting structural properties of objectives under partial information, including hidden convexity, bandit feedback, and preference feedback.
For a complete list of my publications, see below or my Google Scholar.
Here is my CV for more information.
Publications
Keywords: policy gradient methods, general utility RL, convex RL, LLM post-training.
Keywords: learning in games, game dynamics, multi-agent RL, Markov games with general utilities, Markov potential games, continuous games with state dynamics, game theory.
Keywords: online learning, non-convex optimization, hidden convexity, bandit feedback.
Keywords: Stochastic approximation, Dynamical systems, non-convex stochastic optimization, adaptive gradient methods.
SUTD (2024-2025):
ETH Zurich (2022-2024):
Télécom Paris (2018-2021): Teaching Assistant
Optimization for Machine Learning (graduate level), Statistics (graduate level), Probabilities (undergraduate level).