Anas Barakat

Anas Barakat

Research Fellow

Singapore University of Technology and Design

Optimization · RL · Multi-Agent Learning

My research lies at the intersection of optimization, reinforcement learning, and learning in games. I develop foundations and algorithms for sequential decision-making, with a focus on richer objectives, multi-agent interactions, and limited feedback, motivated by emerging challenges in modern AI systems.

I am currently a Research Fellow at Singapore University of Technology and Design, working with Antonios Varvitsiotis. Previously, I was a postdoctoral fellow at ETH Zurich working with Niao He. I received my PhD in mathematics and computer science from Institut Polytechnique de Paris (Télécom Paris), advised by Pascal Bianchi and Walid Hachem.

Optimization for Sequential Decision-Making

I develop optimization foundations and algorithms for sequential decision-making with richer objectives, strategic interactions, and limited feedback. My research is organized around three complementary directions:

  • Policy Optimization Beyond Expected Additive Rewards: Foundations and optimization algorithms for reinforcement learning with objectives beyond standard expected additive rewards, including general utilities and LLM post-training.

  • Learning in Structured Games: Optimization and learning dynamics in strategic and stateful multi-agent environments, including Markov games and continuous games.

  • Online Learning with Time-Varying Objectives and Limited Feedback: Learning and optimization when objectives evolve over time and are only partially observed, including bandit and pairwise comparison feedback.

Selected & Recent Research

On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
Anas Barakat, Souradip Chakraborty, Peihong Yu, Pratap Tokekar, Amrit Singh Bedi
NeurIPS 2025 · Proceedings
Convex Markov Games and Beyond: New Proof of Existence, Characterization and Learning Algorithms for Nash Equilibria
Anas Barakat, Ioannis Panageas, Antonios Varvitsiotis
AISTATS 2026 · Proceedings
Reinforcement Learning with General Utilities: Simpler Variance Reduction and Large State-Action Space
Anas Barakat, Ilyas Fatkhullin, Niao He
ICML 2023 · Proceedings
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
Anas Barakat, Souradip Chakraborty, Khushbu Pahwa, Amrit Singh Bedi
Preprint · arXiv
Online Learning on Hidden-Convex Losses via Algorithmic Equivalence: Optimal Regret, Geometric Barrier, and Bandit Feedback
Anas Barakat, Andreas Kontogiannis, Vasilis Pollatos, Ioannis Panageas, Antonios Varvitsiotis
Preprint · arXiv

For a complete list of my publications, see below or my Google Scholar.

News

07/2026
Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning accepted to TMLR.
05/2026
New preprint: Online Learning on Hidden-Convex Losses via Algorithmic Equivalence: Optimal Regret, Geometric Barrier, and Bandit Feedback. arXiv
05/2026
Recognized as a Gold Reviewer at ICML 2026.
02/2026
New preprint: Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training. arXiv
02/2026
Optimistic Online Learning in Symmetric Cone Games accepted to TMLR. Paper
01/2026
Convex Markov Games and Beyond accepted to AISTATS 2026. arXiv
10/2025
Designed and taught a new course with John Lazarsfeld and Iosif Sakos: Online Learning and Learning in Games.
09/2025
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning accepted to NeurIPS 2025.

Experience and Background

2024–present
Research Fellow
Singapore University of Technology and Design
Working with Antonios Varvitsiotis.
2022–2024
Postdoctoral Fellow
ETH Zürich, Department of Computer Science
Worked with Niao He.
2018–2021
Ph.D. in Applied Mathematics and Computer Science
Institut Polytechnique de Paris (Télécom Paris)
Advised by Pascal Bianchi and Walid Hachem.

Here is my CV for more information.

Learning in Structured Dynamical Games

Keywords: learning in games, game dynamics, multi-agent RL, Markov games with general utilities, Markov potential games, continuous games with state dynamics, game theory.

Online Learning under Limited Feedback

Keywords: online learning, non-convex optimization, hidden convexity, bandit feedback.

Talks

Invited talk - 5th Symposium on Machine Learning and Dynamical Systems
Invited talk - ICCOPT 2025
Invited talk - Learning Theory and Applications Workshop, NTU
Invited talk - Finance and RL Talks
Invited talk - 4th Symposium on Machine Learning and Dynamical Systems

Reviewing

  • Conferences: NeurIPS (‘22,‘23,‘25,‘26), ICML (‘25, Gold Reviewer ‘26), ICLR (‘25,‘26), AISTATS (Top Reviewer ‘22,‘24,‘25,‘26), EC (‘26).

  • Journals: Journal of Machine Learning Research (JMLR), Transactions on Machine Learning Research (TMLR), Mathematical Programming, SIAM Journal on Optimization (SIOPT), Journal of Optimization Theory and Applications (JOTA), IEEE Transactions on Automatic Control.

Teaching

SUTD (2024-2025):

ETH Zurich (2022-2024):

Télécom Paris (2018-2021): Teaching Assistant

Optimization for Machine Learning (graduate level), Statistics (graduate level), Probabilities (undergraduate level).