Skip to main content

Showing 1–4 of 4 results for author: Alakuijala, M

Searching in archive cs. Search in all archives.
.
  1. arXiv:2405.19988  [pdf, other

    cs.RO cs.AI cs.CL cs.CV cs.LG

    Video-Language Critic: Transferable Reward Functions for Language-Conditioned Robotics

    Authors: Minttu Alakuijala, Reginald McLean, Isaac Woungang, Nariman Farsad, Samuel Kaski, Pekka Marttinen, Kai Yuan

    Abstract: Natural language is often the easiest and most convenient modality for humans to specify tasks for robots. However, learning to ground language to behavior typically requires impractical amounts of diverse, language-annotated demonstrations collected on each target robot. In this work, we aim to separate the problem of what to accomplish from how to accomplish it, as the former can benefit from su… ▽ More

    Submitted 30 May, 2024; originally announced May 2024.

    Comments: 10 pages in the main text, 16 pages including references and supplementary materials. 4 figures and 3 tables in the main text, 1 table in supplementary materials

  2. arXiv:2405.15383  [pdf, other

    cs.AI

    Generating Code World Models with Large Language Models Guided by Monte Carlo Tree Search

    Authors: Nicola Dainese, Matteo Merler, Minttu Alakuijala, Pekka Marttinen

    Abstract: In this work we consider Code World Models, world models generated by a Large Language Model (LLM) in the form of Python code for model-based Reinforcement Learning (RL). Calling code instead of LLMs for planning has the advantages of being precise, reliable, interpretable, and extremely efficient. However, writing appropriate Code World Models requires the ability to understand complex instructio… ▽ More

    Submitted 24 May, 2024; originally announced May 2024.

    Comments: 10 pages in main text, 24 pages including references and supplementary materials. 2 figures and 3 tables in the main text, 9 figures and 12 tables when including the supplementary materials

  3. arXiv:2211.09019  [pdf, other

    cs.RO cs.AI cs.CV cs.LG

    Learning Reward Functions for Robotic Manipulation by Observing Humans

    Authors: Minttu Alakuijala, Gabriel Dulac-Arnold, Julien Mairal, Jean Ponce, Cordelia Schmid

    Abstract: Observing a human demonstrator manipulate objects provides a rich, scalable and inexpensive source of data for learning robotic policies. However, transferring skills from human videos to a robotic manipulator poses several challenges, not least a difference in action and observation spaces. In this work, we use unlabeled videos of humans solving a wide range of manipulation tasks to learn a task-… ▽ More

    Submitted 7 March, 2023; v1 submitted 16 November, 2022; originally announced November 2022.

  4. arXiv:2106.08050  [pdf, other

    cs.LG

    Residual Reinforcement Learning from Demonstrations

    Authors: Minttu Alakuijala, Gabriel Dulac-Arnold, Julien Mairal, Jean Ponce, Cordelia Schmid

    Abstract: Residual reinforcement learning (RL) has been proposed as a way to solve challenging robotic tasks by adapting control actions from a conventional feedback controller to maximize a reward signal. We extend the residual formulation to learn from visual inputs and sparse rewards using demonstrations. Learning from images, proprioceptive inputs and a sparse task-completion reward relaxes the requirem… ▽ More

    Submitted 15 June, 2021; originally announced June 2021.