Please use this identifier to cite or link to this item:
https://www.um.edu.mt/library/oar/handle/123456789/148974| Title: | InvPatch : prefix-based conditional generation for inverse dynamics |
| Authors: | Rašajski, Nemanja Makantasis, Konstantinos Liapis, Antonios Yannakakis, Georgios |
| Keywords: | Artificial intelligence Computer vision Machine learning Image processing -- Digital techniques |
| Issue Date: | 2026 |
| Publisher: | Institute of Electrical and Electronics Engineers |
| Citation: | Rašajski, N., Makantasis, K., Liapis, A. & Yannakakis, G. (2026). InvPatch : prefix-based conditional generation for inverse dynamics. 2026 IEEE Conference on Artificial Intelligence (CAI), Granada. 999-1006. |
| Abstract: | Inverse Dynamics Models (IDMs) attempt to predict the actions that cause observable changes in a scene. Current methods for building accurate IDMs either rely on rule-based systems that exploit video metadata or require large-scale training over thousands of video hours. These approaches are inevitably limited to a single domain due to the difficulty of acquiring metadata or sufficient training data across different domains. In response to these challenges, this study draws inspiration from data-efficient video captioning methods, specifically prefix-based conditional generation. This approach maps visual features into prefix tokens that condition the action-prediction process. We introduce InvPatch, a framework that builds on prefix-based conditional generation and extends it by adding learned visual-representation compression. In InvPatch, attention-based patch selection and pooling are applied to features extracted from a ViT backbone, reducing the conditioning input from a set of frame-by-frame features to a single vector. We evaluate our framework across two diverse settings: 3D third-person real world (KIT Bimanual Actions) and 2D synthetic (MUGEN). Our method achieves 97.21% accuracy on MUGEN and surpasses state-of-the-art results on KIT Bimanual Actions. Our sensitivity analysis highlights the data efficiency of this approach, as InvPatch maintains comparable performance even when trained with 30% less data. Additionally, our runtime analysis demonstrates the computational efficiency of InvPatch, which requires fewer trainable parameters and performs fewer FLOPs compared to other methods. |
| URI: | https://www.um.edu.mt/library/oar/handle/123456789/148974 |
| Appears in Collections: | Scholarly Works - FacICTAI |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| invpatch_prefix-based_conditional_generation_for_inverse_dynamics.pdf | 559.48 kB | Adobe PDF | View/Open |
Items in OAR@UM are protected by copyright, with all rights reserved, unless otherwise indicated.
