Skander Moalla

Hi! I’m a final-year PhD candidate at EPFL, advised by Prof. Caglar Gulcehre.

deadline-profile-v3.jpg

My research focuses on scalable learning from experience, driven by three core ingredients: plasticity for continual learning, off-policy/offline reinforcement learning for efficiency, and diversity (parallel) and agentic systems (sequential) for exploration and test-time scaling.

My work includes co-leading the post-training of an open-source 70B LLM (Apertus, ACL’26), leveraging diversity to improve test-time scaling and exploration in RL (Google DeepMind Intern, Gemma post-training), developing an agentic harness for scaling AI theorem proving in Lean (Meta FAIR Intern, MathSI), deriving an offline RL algorithm for LLMs (QRPO, NeurIPS’25), and exposing the connection between plasticity, trust-region, and off-policy collapse (No Rep, No Trust, NeurIPS’24).

I have developed several codebases from scratch with significant infrastructure work such as a scalable code execution sandbox for QRPO, a scalable post-training pipeline for Apertus 70B, and a reproducible ML template. I am also proud of conducting award-winning reproducibility studies (ReScience 2023).

Research

Reinforcement Learning LLM Post-Training, Off-policy RL, Exploration

Test-Time Scaling Diversity, Agentic systems, Formal math

Deep Learning Plasticity, Deep RL, Transformers

Education

PhD in CS EPFL | Prof. Caglar Gulcehre

MSc in CS Oxford | Prof. Shimon Whiteson

BSc in Maths & CS École Polytechnique

Exchange Programs U of Toronto | Software Eng
Stanford | AI & Entrepreneurship

Experience

Research Scientist Intern Meta FAIR (J. Kempe, R. Munos)

PhD Student Researcher Google DeepMind (A. Ramé)

RL Applied Scientist Quincus

SDE/SWE & ML Amazon

Research Assistant Prof. Hess & Blue Brain Project

I am considering full-time positions for winter 2026/2027. I bring a strong background in both research and engineering.

Background

I was a Research Scientist Intern at Meta FAIR (Paris) hosted by Julia Kempe and Rémi Munos in the Foundations of Reasoning / Math Superintelligence Team, working on agentic in-context adaptation, self-improvement, and asymmetric self-play in formal math (Lean). I was also a PhD student researcher at Google DeepMind (Paris) with the Gemma post-training team hosted by Alexandre Ramé, working on diversity to improve test-time scaling and sampling in reinforcement learning. I also interned as an applied scientist at Quincus to optimize middle-mile logistics using RL, and as a software engineer at Amazon building AI tools for Amazon Transportation Services.

I hold an MSc in Advanced Computer Science from the University of Oxford (Reuben College inaugural cohort 2021-2022). My MSc thesis on Multi-Agent RL and policy gradient methods was supervised by Mingfei Sun and Prof. Shimon Whiteson at the Whiteson Research Lab (WhiRL).

I completed a BSc in Mathematics and Computer Science from École Polytechnique (inaugural cohort 2017-2020) and was a visiting student at EPFL, the University of Toronto, and Stanford University, where I focused on software engineering and entrepreneurship.

At EPFL, I was the president of CUBAliente (2024), the Latin social dancing association. At Oxford, I led the Careers team of the Oxford Artificial Intelligence Society (OxAI) (2022) and played as a libero for the men’s Blues volleyball team (2021-2022). At École Polytechnique, I was elected to the first board of the L’ORE (2018), the Bachelor’s program student society.

Selected Publications

  1. ACL 2026
    Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
    Project Apertus,  [...],  Skander Moalla*,  [...], Antoine Bosselut, Martin Jaggi, and Imanol Schlag
    In Annual Meeting of the Association for Computational Linguistics (ACL) 2026
  2. NeurIPS 2025
    Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
    Simon Matrenok*,  Skander Moalla*, and Caglar Gulcehre
    In Advances in Neural Information Processing Systems (NeurIPS) 2025
  3. NeurIPS 2024
    No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO
    Skander Moalla*, Andrea Miele, Razvan Pascanu, and Caglar Gulcehre
    In Advances in Neural Information Processing Systems (NeurIPS) 2024
  4. NeurIPS 2024
    Building on Efficient Foundations: Effectively Training LLMs with Structured Feedforward Layers
    Xiuying Wei,  Skander Moalla, Razvan Pascanu, and Caglar Gulcehre
    In Advances in Neural Information Processing Systems (NeurIPS) 2024
  5. NeurIPS 2023
    SMACv2: An Improved Benchmark for Cooperative Multi-Agent Reinforcement Learning
    Benjamin Ellis, Jonathan Cook,  Skander Moalla, Mikayel Samvelyan, Mingfei Sun, Anuj Mahajan, Jakob Foerster, and Shimon Whiteson
    In Advances in Neural Information Processing Systems (NeurIPS) 2023
  6. ReScience
    [Re] Reproducibility Study of Behavior Transformers
    Skander Moalla*, Manuel Madeira*, Lorenzo Riccio*, and Joonhyung Lee*
    ReScience C 2023