Journal / Conference Paper View Paper

Bari, Italy

LLM-Think-Alouds: Multimodal Player Experience Analysis for VR Playtesting

An IEEE ISMAR 2026 / IEEE TVCG paper on automated multimodal analysis of VR playtesting sessions, using synchronized gameplay video, think-aloud speech, and tone-related cues to estimate perceived difficulty and emotion over time.

Figure from the LLM-Think-Alouds paper showing multimodal VR playtesting analysis

Publication details

  • Authors Erdem Murat, Yongqi Zhang, Liuchuan Yu, Siraj Sabah, and Lap-Fai Yu
  • Publication venue IEEE ISMAR 2026 / IEEE TVCG (conditionally accepted)
  • Location Bari, Italy
  • Date 2026
  • Paper View paper (PDF)

Abstract

Virtual reality (VR) games can evoke strong and dynamic player responses due to their immersive nature. To design and refine these experiences, developers often rely on playtesting to understand how players' emotions and perceived difficulty evolve throughout gameplay. However, manually analyzing gameplay footage and think-aloud data is time-consuming and difficult to scale. We propose a large language model (LLM)-based approach for automated player experience analysis in VR games using synchronized gameplay video and player audio. Based on estimated rating trajectories, the approach supports gameplay analysis, comparison of player experiences, and deeper inspection of how experience unfolds over time. We evaluated the approach through a user study with three VR platformer games. Our findings suggest that LLMs can recover meaningful experiential signals, particularly perceived difficulty, while performance across emotional dimensions varies depending on game context and data expressiveness. Overall, the results demonstrate the potential of LLM-assisted methods as scalable tools for VR playtesting and game design analysis.

Highlights

  • Conditionally accepted to IEEE ISMAR 2026 and IEEE TVCG.
  • Introduces a multimodal foundation-model pipeline for analyzing VR think-aloud playtesting sessions.
  • Estimates and visualizes perceived difficulty and emotion trajectories from synchronized gameplay, speech, and tone cues.