Quality Over Quantity: Image Curation for Multimodal Fake News Detection
编号:37
访问权限:仅限参会人
更新:2026-07-22 16:09:17 浏览:0次
Online
摘要
Most multimodal fake news detection research assumes
that more training images lead to better performance.
We test this assumption directly. Using a Light Cross-Modal
Attention model that fuses frozen DeBERTa-v3-base and CLIP
ViT-B/32 embeddings through a 260K-parameter attention module,
we compare detection accuracy across three image collection
strategies: generic web scraping via Bing (679 images), curated
article-specific images from FakeNewsNet (60 images), and manually
collected images from 15 fact-checking organizations (150
images). In a size-controlled comparison (150 vs. 150), curated
images outperform generic ones by 14.7 percentage points (93.3%
vs. 78.6%, 5-fold CV). The 150 curated images also beat all 679
generic images by 4.5 points. A second experiment on mixedsource
data reveals that multimodal fusion degrades by only 1.4%
when data sources are combined, while text-only and image-only
models drop by 19.6% and 13.5% respectively. These results
point to image-text relevance as a more import
关键词
Terms—fake news detection, multimodal learning, crossmodal attention, image quality, CLIP, DeBERTa
稿件作者
manish rai
Manipal university jaipur
Amrit Raj
Manipal university jaipur
发表评论