Entity Aware Knowledge Guided Vision Language Framework for Image Captioning
编号:42
访问权限:仅限参会人
更新:2026-07-22 16:09:21 浏览:0次
Online
摘要
Entity aware image captioning is an important task which describe the image with the named entity information like person, location, event and organization names. Existing image captioning system generates fluent but generic captions which lacks the informativeness which is purpose for the many downstream tasks. To address this issue, the proposed system proposes hybrid knowledge graph guided Entity aware vision language framework. The proposed method first extracts the named entities and noun phrases that generate target entity memory and support memory which is enhanced using hybrid knowledge graph to further use as input to caption generating decoder with image features to generate the multiple candidate captions. The multiple candidate captions are post generation reranked and verified based on knowledge graph to generate a target caption. The proposed model achieves the CIDEr score of 74.31 and entity F1 of 28.42. Ablation study on selector, reranker and verifier shows that the verifier-based safety control has reduced the unsupported entities while forming entity aware caption.
关键词
Knowledge Graph,Entity Awareness,Reranking,Hallucination Mitigation,Context Awareness,Diversity in image captioning
稿件作者
Sharmila Kharat
Assistant Professor
Sunita Barve
MIT Academy of Engineering Alandi Pune
发表评论