Abstractive Summarization of Bengali Academic Videos Based on Audio Subtitles

Anyone who has studied online knows the feeling. You need one specific explanation from a two-hour lecture, and you end up dragging the progress bar back and forth for twenty minutes. Or you missed a live class because of a schedule clash or a time zone difference, and now you have a recording you do not have time to watch in full.

For English, there are tools that help. For Bengali, a language spoken by more than 300 million people and widely used in education in Bangladesh and West Bengal, there was essentially nothing. In our paper, Abstractive Summarization of Bengali Academic Videos Based on Audio Subtitles, published in the Findings of EACL 2026, we built the first complete pipeline that takes a Bengali academic video and gives you a summary, a title, and timestamps.

This one is a bit different from my usual multimodal work. It is a collaboration led by students at KUET in Bangladesh, together with researchers from HBKU, UC Riverside and UMass Lowell, and I am happy I could be part of it.

Why Is This Hard?

Summarizing a video sounds like one task, but it is really several problems stacked on top of each other:

  1. Speech recognition. You first need to turn spoken Bengali into text. Existing Bengali speech recognition systems are not very accurate, and any mistake here damages everything that follows.

  2. Informal, spoken language. Most Bengali summarization models were built on news articles, which are short, formal and carefully edited. A lecture is the opposite: conversational, repetitive, and full of half-finished sentences.

  3. Long inputs. A lecture transcript is far longer than what a language model can read at once.

  4. Titles and navigation. A useful tool should also give the video a sensible title and tell you where in the video each part of the summary comes from.

Existing Bengali datasets did not cover any of this. They were made for formal text, so models trained on them do not handle spoken academic content well. In our experiments, that gap turned out to be huge (more on that below).

Our Solution: A Complete Pipeline

The system goes from raw video to a finished summary in a handful of steps.

  • Step 1: Extract and clean up the audio. We pull the audio track out of the video and split it into 45-second chunks. Lecture audio can be quiet, so we measure the loudness of each chunk and boost it by a small amount if needed (up to 8 dB for the quietest chunks). We only normalize if the loudest peaks get too close to distortion. Doing this per chunk is efficient and avoids the distortion that a single global adjustment can cause.

  • Step 2: Transcribe the speech. We tested seven speech recognition approaches on Bengali, including wav2vec2, Facebook MMS and several Whisper variants. Google’s Universal Speech Model, accessed through its speech API, was the clear winner. Its word error rate was 0.3672, and after our audio pre-processing it dropped to 0.2825, about a 23% improvement. The Whisper models, by comparison, had error rates above 1.0 on this data.

  • Step 3: Add punctuation. The transcript comes out with no punctuation, which makes it hard for a model to find sentence boundaries. We insert punctuation using an open-source Bengali punctuation model so the text splits into proper sentences.

  • Step 4: Smart chunking. The summarization model, BanglaT5, can only read 512 tokens at a time, and a lecture is much longer. Instead of cutting at arbitrary points, we split the transcript into chunks made of whole sentences, and each chunk overlaps with the next by one sentence so context is not lost at the boundaries.

    We compared this with two alternatives. On ROUGE-1, overlapping chunks scored 0.5360, non-overlapping chunks 0.4930, and a “progressive” method that summarizes chunk by chunk and re-summarizes 0.3480. It is also cheaper than the progressive method because it avoids repeated processing.

  • Step 5: Summarize and title.

    • For summarization, we fine-tuned BanglaT5 on our new dataset of 10,029 text-summary pairs. Each chunk is summarized with beam search. For very long videos with more than 50 chunks, we summarize recursively: merge neighboring summaries and summarize again until the count is small enough.
    • For title generation, we fine-tuned mT5-multilingual-XLSum on 1,005 summary-title pairs to write a short title from the summary.
  • Step 6: Timestamps and clean-up. Since we know the start time of every 45-second audio chunk, we spread the time across the sentences in each chunk to give every part of the summary a start and end time. That is what lets you jump to the right moment in the video. A final post-processing step also converts spoken forms into standard notation, for example spelled-out numbers into digits, and spoken chemical names, units and math symbols into their usual written forms (think “sin θ”, “HCl” or “kg”).

Everything ran on a single NVIDIA T4 GPU on Google Colab, so it does not need a big cluster.

The Two New Datasets

Since no suitable data existed, we made our own, and both are released with the paper.

  • Summarization dataset: 10,029 text-summary pairs from 213 Bengali academic videos covering 6 subjects and 46 topics. The summaries were written by hand. More than 200 different speakers appear in the videos, which brings in a range of dialects and speaking styles.

  • Title generation dataset: 1,005 summary-title pairs from 335 videos across 54 topics.

Both were split 80% for training, 10% for validation and 10% for testing. The videos come from public YouTube playlists, used with the content owners’ consent and only for academic research.

What We Found

Summarization. We compared four fine-tuned models: BanglaT5, mBART-50, NLLB-200-Distilled and mT5 (small). BanglaT5 came out on top, with a BERTScore F1 of 0.8793, ROUGE-1 of 0.3894 and ROUGE-L of 0.2557 in our main benchmark. It beat the second-best model by about 10.7% on ROUGE-1, 38.6% on ROUGE-2 and 14.7% on ROUGE-L.

Full videos versus recent general-purpose models. We also tried several recent models (LLaVA, Qwen2.5-1.5B, Phi-3-mini and Llama 7B) on full-video summaries without fine-tuning. Most scored close to zero on ROUGE, and only Llama did moderately well (ROUGE-1 of 0.277). The fine-tuned BanglaT5 reached 0.528. For a low-resource language and a specialized domain, fine-tuning on the right data clearly matters.

Human evaluation. Seven independent reviewers, who did not know whether a summary was human-written or model-generated, rated 100 samples on grammar, information coverage, factual consistency and conciseness. The model’s summaries scored close to the human-written ones.

Titles. Out of the box, none of the models could write good Bengali academic titles. After fine-tuning, mT5-multilingual-XLSum reached ROUGE-1 and ROUGE-L F1 scores of 0.4476 and 0.3720, well ahead of BanglaT5 on this task.

Does formal text data help? We fine-tuned BanglaT5 on two popular formal Bengali summarization datasets and tested it on our spoken academic data. The ROUGE-1 scores were 0.063 and 0.190, compared with 0.389 when training on our dataset. That is the clearest evidence that spoken content needs its own data.

Why This Matters

The practical goal is accessibility. Students who cannot attend live sessions, who have limited time, or who just want to find one topic in a long recording can get a quick overview and jump to the right moment. Because the system also shows strong zero-shot behavior on other kinds of spoken Bengali content, it can serve as a solid baseline for future work in the language.

More broadly, a lot of the AI progress we hear about happens in English. Building for languages like Bengali takes careful data work, not just bigger models.

Limitations and Future Work

We worked with text summaries derived from the subtitles, not summaries in video form. A natural next step is multimodal summarization that combines visual, audio and text signals to produce short video clips. Bengali dialects are also very diverse, and even with 200+ speakers we need broader representation. Transcription errors still propagate into the summaries, so better Bengali speech recognition would lift the whole pipeline. Finally, real-time summarization of live sessions is another promising direction, but it needs efficient optimization to process continuous audio with low delay.

If you are interested in my main line of research on combining different kinds of data such as images, audio and text, take a look at my posts on Cross-Modal Proxy Tokens and MMSFormer .

More Details About the Paper

comments powered by Disqus

Related Posts

Hello World, I Am Back!

Hello world, I am back! If you have visited this blog before, you may have noticed that things went quiet for a while.

Read more

MSP: Modality Self-Play for Training Multimodal LLMs Without Paired Cross-Modal Data

Think about how you understand a video of a dog barking in a park. Your eyes see the dog, your ears hear the bark, and your brain effortlessly puts the two together.

Read more

DualSwinFusionSeg: Multimodal Martian Landslide Segmentation via Dual Swin Transformer with Multi-Scale Fusion and UNet++

Mars has landslides. Just like on Earth, slopes collapse, dust and rock slide downhill, and the scars stay visible for a long time.

Read more