MIRMI Aktuelles
From Video Understanding to Embodied Agents - Munich AI Lectures
VERANSTALTUNGEN |
From Video Understanding to Embodied Agents
Ivan Laptev (MBZUAI) | Munich AI Lectures | 25. June 2024, 17:00 – 18:00 CET
Location: Arcisstr. 21, 80333 Munich, Room 0790 (ground floor)

Ivan Laptev
To the lecture series
More info on Prof. Laptev's work: https://mbzuai.ac.ae/study/faculty/ivan-laptev/
Ivan Laptev is a visiting professor at MBZUAI and a senior researcher on leave from Inria Paris. He received a PhD degree in Computer Science from the Royal Institute of Technology in 2004 and a Habilitation degree from École Normale Supérieure in 2013.
Ivan’s main research interests include visual recognition of human actions, objects and interactions, and more recently robotics. He has published over 120 technical papers most of which appeared in international journals and major peer-reviewed conferences of the field. He served as an associate editor of IJCV and TPAMI, he served as a program chair for CVPR’18 and ICCV’23, he will serve as a program chair for ACCV’24 and a general chair for ICCV’29. He has co-organized several tutorials, workshops and challenges at major computer vision conferences. He has also co-organized a series of INRIA summer schools on computer vision and machine learning (2010-2013) and Machines Can See summits (2017-2024). He received an ERC Starting Grant in 2012 and was awarded a Helmholtz prize for significant impact on computer vision in 2017.
About the talk
Computer vision has recently excelled on a wide range of tasks such as image classification, segmentation and captioning. This impressive progress now powers many applications of internet imaging and yet, current methods still fall short in addressing embodied understanding of visual scenes. What will happen if pushing an object over a table border? What precise actions are required to plant a tree? Building systems that can answer such questions from visual inputs will empower future applications of robotics and personal visual assistants while enabling methods to operate in unstructured real-world environments.
Following this motivation, in this talk we will address models and learning methods that derive procedural knowledge from instructional videos. I will then describe our recent work on visual manipulation and will present a new dataset for long-term story-level video understanding.
---
On a monthly basis, Munich AI Lectures invite top-level AI researchers to give a glimpse into their work and the future of AI. The Munich AI Lectures are a joint initiative of the baiosphere, Bavarian Academy of Science and Humanities (BAdW), Helmholtz Munich, Ludwig Maximilian University of Munich (LMU), Technical University of Munich (TUM), AI-HUB@LMU, ELLIS Chapter Munich, Konrad Zuse School of Excellence in Reliable AI (relAI), Munich Center for Machine Learning (MCML), Munich Data Science Institute (MDSI) at TUM, and Munich Institute of Robotics and Machine Intelligence (MIRMI).
The lectures consist of a short presentation followed by a Q&A to enable a lively discussion with our speakers. Each lecture lasts about one hour and will be streamed live on Munich AI Lectures' YouTube Channel. Recordings will be available afterwards.