My research develops foundations and algorithms for sequential and strategic decision-making when objectives extend beyond expected cumulative reward, feedback is limited, and multiple adaptive agents interact. My work lies at the intersection of optimization, reinforcement learning, and learning in games. I am particularly motivated by emerging challenges in modern AI, including policy optimization for LLM post-training, preference-based learning, and decentralized learning in competitive and cooperative multi-agent systems.
I am currently a Research Fellow at Singapore University of Technology and Design, working with Georgios Piliouras and Antonios Varvitsiotis. Previously, I was a postdoctoral fellow at ETH Zurich working with Niao He. I received my PhD in mathematics and computer science from Institut Polytechnique de Paris (Télécom Paris), advised by Pascal Bianchi and Walid Hachem.
Optimization for Sequential and Strategic Decision-Making
My research develops foundations and algorithms for decision-making beyond classical assumptions on objectives, feedback, and agent interactions, spanning three complementary directions:
Policy optimization beyond expected additive rewards: Foundations and optimization algorithms for reinforcement learning with objectives beyond standard expected additive rewards, including general utilities and LLM post-training objectives.
Learning and optimization in structured games: Foundations and algorithms for learning in strategic and dynamic multi-agent environments, with a focus on the long-run behavior of learning dynamics and convergence to equilibria.
Online learning beyond convexity under limited feedback: Understanding when learning remains possible with nonconvex objectives and only partial or ordinal feedback. I develop algorithms and guarantees by exploiting latent structure—such as hidden convexity—and learning directly from bandit or preference feedback.
For a complete list of my publications, see below or my Google Scholar.
Here is my CV for more information.
Publications
Keywords: policy gradient methods, general utility RL, convex RL, LLM post-training.
Keywords: learning in games, game dynamics, multi-agent RL, Markov games with general utilities, Markov potential games, continuous games with state dynamics, game theory.
Keywords: online learning, bandit feedback, preference feedback, hidden convexity, non-convex optimization.
Keywords: Stochastic approximation, Dynamical systems, non-convex stochastic optimization, adaptive gradient methods.
Olivier Lepel — Master’s thesis, Mar.–Sep. 2024
Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning. TMLR 2026.
Currently: Data Scientist at BCG.
Kimon Protopapas — Master’s semester project, Sep. 2023–Jan. 2024
Policy Mirror Descent with Lookahead. NeurIPS 2024.
Currently: Research Intern at University of Basel.
Philip Jordan — Master’s thesis, May–Dec. 2023
Independent Learning in Constrained Markov Potential Games. AISTATS 2024.
Currently: PhD student at EPFL.
Jiduan Wu — Master’s thesis, Oct. 2022–Mar. 2023
Learning Zero-Sum Linear Quadratic Games with Improved Sample Complexity and Last-Iterate Convergence. IEEE Conference on Decision and Control (CDC) 2023, SIAM Journal on Control and Optimization, 2025.
Currently: PhD student at Max Planck ETH Center for Learning Systems.
Harish Rajagopal — Master’s thesis, Mar.–Sep. 2022
Multistage Step Size Scheduling for Minimax Problems.
Currently: Quantitative Technologist at Qube Research & Technologies.
Julien Lehmann — Master’s thesis, May–Oct. 2021
Analysis of a Target-Based Actor-Critic Algorithm with Linear Function Approximation. AISTATS 2022.
Currently: Research Engineer at OCamlPro.
SUTD (2024-2025):
ETH Zurich (2022-2024):
Télécom Paris (2018-2021): Teaching Assistant
Optimization for Machine Learning (graduate level), Statistics (graduate level), Probabilities (undergraduate level).