My research develops foundations and algorithms for sequential and strategic decision-making when objectives extend beyond expected cumulative reward, feedback is limited, and multiple adaptive agents interact. My work lies at the intersection of optimization, reinforcement learning, and learning in games. I am particularly motivated by emerging challenges in modern AI, including policy optimization for LLM post-training, preference-based learning, and decentralized learning in competitive and cooperative multi-agent systems.
I am currently a Research Fellow at Singapore University of Technology and Design, working with Georgios Piliouras and Antonios Varvitsiotis. Previously, I was a postdoctoral fellow at ETH Zurich working with Niao He. I received my PhD in mathematics and computer science from Institut Polytechnique de Paris (Télécom Paris), advised by Pascal Bianchi and Walid Hachem.
Optimization for Sequential and Strategic Decision-Making
My research develops foundations and algorithms for decision-making beyond classical assumptions on objectives, feedback, and agent interactions, spanning three complementary directions:
Policy Optimization Beyond Expected Additive Rewards: Foundations and optimization algorithms for reinforcement learning with objectives beyond standard expected additive rewards, including general utilities and LLM post-training.
Learning and Optimization in Structured Games: Foundations and algorithms for strategic and stateful multi-agent environments, with a focus on equilibria and learning dynamics.
Online Learning with Structured Objectives under Limited Feedback: Learning and optimization by exploiting structural properties of objectives under partial information, including hidden convexity, bandit feedback, and preference feedback.
For a complete list of my publications, see below or my Google Scholar.
Here is my CV for more information.
Publications
Keywords: policy gradient methods, general utility RL, convex RL, LLM post-training.
Keywords: learning in games, game dynamics, multi-agent RL, Markov games with general utilities, Markov potential games, continuous games with state dynamics, game theory.
Keywords: online learning, non-convex optimization, hidden convexity, bandit feedback.
Keywords: Stochastic approximation, Dynamical systems, non-convex stochastic optimization, adaptive gradient methods.
SUTD (2024-2025):
ETH Zurich (2022-2024):
Télécom Paris (2018-2021): Teaching Assistant
Optimization for Machine Learning (graduate level), Statistics (graduate level), Probabilities (undergraduate level).