Please use this identifier to cite or link to this item: https://www.um.edu.mt/library/oar/handle/123456789/148974
Title: InvPatch : prefix-based conditional generation for inverse dynamics
Authors: Rašajski, Nemanja
Makantasis, Konstantinos
Liapis, Antonios
Yannakakis, Georgios
Keywords: Artificial intelligence
Computer vision
Machine learning
Image processing -- Digital techniques
Issue Date: 2026
Publisher: Institute of Electrical and Electronics Engineers
Citation: Rašajski, N., Makantasis, K., Liapis, A. & Yannakakis, G. (2026). InvPatch : prefix-based conditional generation for inverse dynamics. 2026 IEEE Conference on Artificial Intelligence (CAI), Granada. 999-1006.
Abstract: Inverse Dynamics Models (IDMs) attempt to predict the actions that cause observable changes in a scene. Current methods for building accurate IDMs either rely on rule-based systems that exploit video metadata or require large-scale training over thousands of video hours. These approaches are inevitably limited to a single domain due to the difficulty of acquiring metadata or sufficient training data across different domains. In response to these challenges, this study draws inspiration from data-efficient video captioning methods, specifically prefix-based conditional generation. This approach maps visual features into prefix tokens that condition the action-prediction process. We introduce InvPatch, a framework that builds on prefix-based conditional generation and extends it by adding learned visual-representation compression. In InvPatch, attention-based patch selection and pooling are applied to features extracted from a ViT backbone, reducing the conditioning input from a set of frame-by-frame features to a single vector. We evaluate our framework across two diverse settings: 3D third-person real world (KIT Bimanual Actions) and 2D synthetic (MUGEN). Our method achieves 97.21% accuracy on MUGEN and surpasses state-of-the-art results on KIT Bimanual Actions. Our sensitivity analysis highlights the data efficiency of this approach, as InvPatch maintains comparable performance even when trained with 30% less data. Additionally, our runtime analysis demonstrates the computational efficiency of InvPatch, which requires fewer trainable parameters and performs fewer FLOPs compared to other methods.
URI: https://www.um.edu.mt/library/oar/handle/123456789/148974
Appears in Collections:Scholarly Works - FacICTAI

Files in This Item:
File Description SizeFormat 
invpatch_prefix-based_conditional_generation_for_inverse_dynamics.pdf559.48 kBAdobe PDFView/Open


Items in OAR@UM are protected by copyright, with all rights reserved, unless otherwise indicated.