Skip to main content

VICTOR-XR partners present new AI research at BMVC 2026

Researchers from DFKI (Coordinator of the VICTOR-XR project ) and RPTU Kaiserslautern-Landau will present their research paper “MoTE: Mixture of Task Experts for Multi-Task Video Understanding” at the 37th British Machine Vision Conference (BMVC 2026), taking place from 23 to 26 November 2026 in Lancaster, UK.

BMVC is a major international conference on computer vision, machine learning, and related areas, bringing together researchers working on topics ranging from image and video understanding to multimodal learning and AI.

The paper, authored by Muhammad Asad Ali, Umar Khan, Nadia Robertini and Didier Stricker, presents a new way of designing AI systems that need to perform different tasks.

Giving AI different “experts” for different jobs

Imagine an AI system watching a video. It might be asked to recognise an action, predict what will happen next or understand a sequence of steps. These are different tasks, so why should the AI process all of them in exactly the same way?

MoTE introduces specialised AI “experts” for different tasks. The system keeps a shared core but activates the expert most relevant to the task at hand.

This means that, instead of making the entire model work harder every time a new task is requested, only the relevant specialised part needs to be activated.

Tested on video understanding

The researchers tested MoTE on the COIN benchmark, which evaluates how well AI systems understand instructional and procedural videos.

The MoTE model achieved an average accuracy of 62.9% across five video-understanding tasks, the highest average reported among the methods included in the comparison.

The researchers also tested the same idea beyond video. By adding a specialised expert to an OCR model, they enabled it to extract structured information from receipt images, while keeping its original OCR capability.

Why this matters for VICTOR-XR

The ability to combine shared knowledge with specialised AI capabilities is relevant to developing more flexible multimodal systems.

This is closely connected to VICTOR-XR’s broader vision, which explores combining AI, extended reality, and multimodal interaction to create new ways for people and digital systems to interact.

The MoTE research offers one example of how AI can be made more adaptable by giving different tasks access to the capabilities they actually need.

Explore MoTE

The researchers have also created an interactive demonstration showing how a video and a prompt enter the system, how the appropriate expert is selected and how the AI generates its response.

🎥 Watch the MoTE demonstration

📄 Read the paper: MoTE: Mixture of Task Experts for Multi-Task Video Understanding

💻 Explore the code on GitHub

Funded by the European Union under Grant Agreement Nº 101298864. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Commission. Neither the European Union nor the European Commission can be held responsible for them