Computer Science

Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Thu Nguyen, Jan Held, A. Cioppa, Silvio Giancola, V. Thambawita, Michael A. Riegler, Pål Halvorsen

2025.10.24International Journal of Semantic Computing

DOI: 10.1142/s1793351x25450035

tlooto Summary

The experiments reveal that incorporating ASR-generated transcripts as a third modality alongside audio and video can improve the performance of multimodal event detection, and that feeding powerful LLMs like Gemini-1.5-Pro with visual data may not improve results compared to their text-only counterpart, but rather degrade performance.

Abstract

We present SoccerNet-Echoes, an extension of the SoccerNet dataset which has been curated by augmenting the 550 games in the original dataset with multilingual audio commentary transcriptions, with a pipeline utilizing OpenAI’s Whisper models for transcription and Google Translate for translation to English. We demonstrate the potential of SoccerNet-Echoes through several applications. Our experiments reveal that incorporating ASR-generated transcripts as a third modality alongside audio and video can improve the performance of multimodal event detection, with our audio-video-text model achieving a top F1-score of 0.7175. We also introduce a novel framework that leverages Large Language Models (LLMs) to extract both predefined, official events, as well as unscripted, unofficial events directly from the commentary. Our evaluation shows that the Gemini-1.5-Pro model effectively identifies official events from text alone, and that LLM-generated game summaries are more descriptive and accurate when using SoccerNet-Echoes compared to using only structured event data. Surprisingly, our experiments also show that feeding powerful LLMs like Gemini-1.5-Pro with visual data may not improve results compared to their text-only counterpart, but rather degrade performance, for which we analyze the potential reasons. By releasing SoccerNet-Echoes, we provide a resource for the scientific community and offer benchmarks that highlight the current capabilities and limitations of ASR and LLM technologies in the domain of multimodal sports analysis.

Citation format

SARKHOOSH, Mehdi Houshmand, et al. Beyond audio: Enhancing soccernet-echoes with multimodal event extraction using LLMs. International Journal of Semantic Computing, 2025, 19: 589–613.