My research lies at the intersection of optimization, reinforcement learning, and learning in games. I develop foundations and algorithms for sequential decision-making, with a focus on richer objectives, multi-agent interactions, and limited feedback, motivated by emerging challenges in modern AI systems.
I am currently a Research Fellow at Singapore University of Technology and Design, working with Antonios Varvitsiotis. Previously, I was a postdoctoral fellow at ETH Zurich working with Niao He. I received my PhD in mathematics and computer science from Institut Polytechnique de Paris (Télécom Paris), advised by Pascal Bianchi and Walid Hachem.
Optimization for Sequential Decision-Making
I develop optimization foundations and algorithms for sequential decision-making with richer objectives, strategic interactions, and limited feedback. My research is organized around three complementary directions:
Policy Optimization Beyond Expected Additive Rewards: Foundations and optimization algorithms for reinforcement learning with objectives beyond standard expected additive rewards, including general utilities and LLM post-training.
Learning in Structured Games: Optimization and learning dynamics in strategic and stateful multi-agent environments, including Markov games and continuous games.
Online Learning with Time-Varying Objectives and Limited Feedback: Learning and optimization when objectives evolve over time and are only partially observed, including bandit and pairwise comparison feedback.
For a complete list of my publications, see below or my Google Scholar.
Here is my CV for more information.
Publications
Keywords: policy gradient methods, general utility RL, convex RL, LLM post-training.
Keywords: learning in games, game dynamics, multi-agent RL, Markov games with general utilities, Markov potential games, continuous games with state dynamics, game theory.
Keywords: online learning, non-convex optimization, hidden convexity, bandit feedback.
Keywords: Stochastic approximation, Dynamical systems, non-convex stochastic optimization, adaptive gradient methods.
Conferences: NeurIPS (‘22,‘23,‘25,‘26), ICML (‘25, Gold Reviewer ‘26), ICLR (‘25,‘26), AISTATS (Top Reviewer ‘22,‘24,‘25,‘26), EC (‘26).
Journals: Journal of Machine Learning Research (JMLR), Transactions on Machine Learning Research (TMLR), Mathematical Programming, SIAM Journal on Optimization (SIOPT), Journal of Optimization Theory and Applications (JOTA), IEEE Transactions on Automatic Control.
SUTD (2024-2025):
ETH Zurich (2022-2024):
Télécom Paris (2018-2021): Teaching Assistant
Optimization for Machine Learning (graduate level), Statistics (graduate level), Probabilities (undergraduate level).