Multimodal Methods for Educational Applications


Megha Mariam K. M

Abstract

Recent advances in large language models (LLMs), vision–language models (VLMs), and text-to- video (T2V) systems have sparked growing excitement about the future of education. Imagine a learn- ing experience where complex ideas can be explained through synchronized text, visuals, narration, and video; where educational content can be generated on demand; and where learning becomes more interactive, accessible, and personalized. While these possibilities are increasingly within reach, an im- portant question remains: do current multimodal AI systems truly understand and connect information across different modalities in ways that meaningfully support learning? This thesis explores that ques- tion through the development of benchmarks and evaluation frameworks for multimodal educational AI.

A familiar challenge for learners is following presentation slides while simultaneously listening to a speaker. Slides are often dense, explanations move quickly, and learners can easily lose track of where to focus their attention. To address this, we explore a method that automatically highlights slide regions relevant to the ongoing narration, helping guide learners toward the most important visual information in real time. Supporting this effort, we introduce a dataset designed to evaluate how effectively models can align spoken language with slide content. The dataset consists of 14 presentation videos and 150 slides, where each slide is paired with its corresponding audio segment and transcript. It captures a wide range of visual and textual elements—including figures, equations, tables, and diverse layouts—making it a challenging and realistic benchmark for multimodal alignment.

 

Year of completion:  June 2026
 Advisor :

Prof. C.V. Jawahar


Related Publications


    Downloads

    thesis