Logo EPFL, École polytechnique fédérale de Lausanne

BIOROB

    Afficher / masquer le formulaire de recherche
    Masquer le formulaire de recherche
    • EN
    Menu
    1. Laboratories
    2. Biorobotics Laboratory (BioRob)
    3. Teaching and student projects
    4. Student project list

    Project Database

    This page contains the database of possible research projects for master and bachelor students in the Biorobotics Laboratory (BioRob). Visiting students are also welcome to join BioRob, but it should be noted that no funding is offered for those projects (see https://biorob.epfl.ch/students/ for instructions). To enroll for a project, please directly contact one of the assistants (directly in his/her office, by phone or by mail). Spontaneous propositions for projects are also welcome, if they are related to the research topics of BioRob, see the BioRob Research pages and the results of previous student projects.

    Search filter: only projects matching the keyword Computer Science are shown here. Remove filter

    Amphibious robotics
    Computational Neuroscience
    Dynamical systems
    Human-exoskeleton dynamics and control
    Humanoid robotics
    Miscellaneous
    Mobile robotics
    Modular robotics
    Neuro-muscular modelling
    Quadruped robotics


    Mobile robotics

    773 – World model and RL for robot manipulation
    Show details
    Category:master project (full-time)
    Keywords:Computer Science, Machine learning, Python, Robotics
    Type:45% theory, 5% hardware, 50% software
    Responsibles: (undefined, phone: 37432)
    (MED11626, phone: 41783141830)
    Description:
    This project has been taken for 2026 Fall semester.

    INTRODUCTION

    Generalist Vision-Language-Action (VLA) policies can now perform a wide range of robot manipulation skills, but evaluating and improving them remains slow and expensive: rigorous evaluation requires hundreds of real-world rollouts, and systematic improvement demands additional corrective data with expert labels. Generative world models offer a scalable alternative by letting policies roll out inside imagination space, and recent work has shown that controllable, multi-view, action-conditioned world models can both rank policy performance without real-world execution and boost success rates by synthesising successful trajectories for supervised fine-tuning. In parallel, goal-conditioned formulations specify tasks through a target visual or language goal rather than a scalar reward, enabling reward-free online improvement through hindsight relabelling.

    OBJECTIVES

    This project will tightly couple these three directions on our existing manipulation platforms at EPFL. Concretely, we will:

    1. Integrate a controllable multi-view world model with our ROS2-based robot platforms (ViperX 300S and WidowX-250 arms, optionally a humanoid such as Reaman), including time-synchronisation with RGB-D cameras and robot state.
    2. Adapt the world model into a goal-conditioned predictor that, given a current observation and a visual or language goal, imagines temporally consistent multi-view futures.
    3. Implement an iterative co-improvement loop in which a small batch of real-world policy rollouts is used to ground the world model on contact-rich, deformable-object tasks, and the grounded model in turn generates large-scale synthetic data.
    4. Replace the explicit success-classifier reward model with hindsight goal relabelling and LoRA-based fine-tuning of a π0.5-class VLA policy, enabling reward-free autonomous improvement.
    5. Perform systematic evaluation, ablation, and documentation to deliver a reusable goal-conditioned RL pipeline for future projects.
    IMPORTANCE

    We have well-documented tutorials on using the robots, teleoperation interfaces for data collection, the HPC cluster, and a complete pipeline for training robot policies. Open-source codebases for π0.5 (openpi), Ctrl-World, Act2Goal, and video diffusion will be integrated into this ecosystem, so the student can focus on the research questions rather than low-level setup.

    We already have a strong base and results from an ongoing project: world model and RL for robot manipulation. Compared to standard language-conditioned VLAs trained only on demonstrations, this paradigm gives us:

    • Imagination-based policy evaluation that closely tracks real-world performance rankings, with no physical rollouts.
    • Targeted policy improvement on unseen objects, novel spatial layouts, and out-of-distribution instructions via synthetic successful trajectories.
    • Physically grounded predictions of contact-rich and deformable-object interactions once the world model is fine-tuned on online rollout data.
    • Reward-free online adaptation through hindsight relabelling.
    WHAT WE HAVE
    1. Ready-and-easy-to-use robot platforms: ViperX 300S and WidowX-250 arms configured with 4 RealSense D405 cameras, various grippers, and a mobile robot platform, fully compatible with the LeRobot V3.0 data format.
    2. Pretrained models: access to π0.5 / π0-FAST policies, the Ctrl-World world-model checkpoint, and Stable Video Diffusion backbones.
    3. Computing resources: two desktop PCs with NVIDIA GPUs 5090 and 4090, plus HPC cluster access for world-model fine-tuning.
    CANDIDATES

    Interested students can apply by emailing sichao.liu@epfl.ch or lixuan.tang@epfl.ch. Please attach your transcript and a short description of your past/current experience on related topics such as robotics, computer vision, reinforcement learning, generative models, and VLA.

    The position is open until we have final candidates. Otherwise, the position will be closed.

    RECOMMENDED READING
    1. “Ctrl-World: A Controllable Generative World Model for Robot Manipulation.” ICLR 2026. arXiv:2510.10125.
    2. “VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model.” arXiv preprint arXiv:2602.12063, 2026.
    3. “Act2Goal: From World Model To General Goal-conditioned Policy.” arXiv preprint arXiv:2512.23541, 2025.
    4. “LeRobot: An Open-Source Library for End-to-End Robot Learning.” arXiv preprint arXiv:2602.22818, 2026.
    5. “π0.5: A Vision-Language-Action Model with Open-World Generalization.” arXiv preprint arXiv:2504.16054, 2025.
    6. “Evaluating Robot Policies in a World Model.” arXiv preprint arXiv:2506.00613, 2025.

    Last edited: 13/08/2026

    One project found.

    Quick links

    • Teaching and student projects
      • Project database
      • Past student projects
      • Students FAQ
    Logo EPFL, École polytechnique fédérale de Lausanne
    • Contact
    • Alessandro Crespi
    • +41 21 693 66 30
    Accessibility Disclaimer

    © 2026 EPFL, all rights reserved