Abstract
Product attributes on e-commerce pages are often distributed across two modalities, textual content and embedded images. Such crossmodal presentation can introduce inconsistencies in wording, units, values, and attribute presence. Existing approaches match representations across modalities, but such alignment only indicates correspondence and cannot determine the relation type between paired attributes, such as equivalence and value conflict.We propose a framework that verifies attribute consistency at the attribute-pair level. First, it aligns cross-modal attributes by combining representation normalization with semantic embedding matching. Second, it classifies each aligned pair into four relation types. Third, it represents the inferred relations as typed triples in a relational graph for relational embedding learning. Finally, we evaluate the proposed framework on a real-world dataset in terms of relation classification, link prediction, and structural triple injection.
Highlights