Loading...
Human-centered attention for VQA in autonomous driving: from machine perception to human-aligned understanding
Citations
Altmetric:
Date
2025-12
Abstract
Visual Question Answering (VQA) is a critical task at the intersection of computer vision and NLP. While current VQA models demonstrate strong performance in general-purpose domains, their effectiveness diminishes in high-stakes contexts such as autonomous driving. These models often lack the ability to prioritise semantically relevant features in dynamic environments, leading to reduced interpretability and trustworthiness, two factors essential for certifiable deployment in safety-critical systems.
This thesis addresses these limitations by introducing a human-centred approach to VQA for autonomous driving. Through a comprehensive evaluation of existing VQA models, key performance gaps have been identified when applied to driving-related queries. To better capture the contextual nuances of such scenarios, a Subjective Scoring Framework has been developed that incorporates human judgment into model evaluation, enabling a more rigorous assessment of response relevance.
The misalignment between human and model attention has been further investigated by conducting comparative studies using attention visualisations. This analysis informs the design of a Human Attention Filter, a novel mechanism that integrates gaze data from eye-tracking experiments into the model’s visual processing pipeline. To facilitate this objective, a domain-specific dataset, called DriVQA, has been developed, which integrates driving images, question–answer pairs, and associated human gaze patterns.
Empirical results demonstrate that incorporating human attention into VQA models leads to substantial improvements in both answer quality and visual reasoning alignment. The filtered models exhibit stronger semantic focus, higher response accuracy, and greater interpretability. By aligning model behaviour with human cognitive patterns, this research contributes to the creation of AI systems that are both functionally robust and cognitively interpretable, which is an essential foundation for the responsible deployment of artificial intelligence in autonomous driving contexts.
Supervisor
Description
Peer-reviewed
Publisher
University of Limerick
Citation
Collections
Files
Loading...
Rekanar_2025_Human.pdf
Adobe PDF, 64.42 MB
ULRR Identifiers
Funding code
Funding Information
Sustainable Development Goals
External Link
License
Attribution-NonCommercial-ShareAlike 4.0 International
