Global ETD Search

Return to search

Vyvozování v přirozeném jazyce s využitím obrazových dat / Grounding Natural Language Inference on Images

Grounding Natural Language Inference on Images Hoa Trong VU July 20, 2018 Abstract Despite the surge of research interest in problems involving linguistic and vi- sual information, exploring multimodal data for Natural Language Inference remains unexplored. Natural Language Inference, regarded as the basic step towards Natural Language Understanding, is extremely challenging due to the natural complexity of human languages. However, we believe this issue can be alleviated by using multimodal data. Given an image and its description, our proposed task is to determined whether a natural language hypothesis contra- dicts, entails or is neutral with regards to the image and its description. To address this problem, we develop a multimodal framework based on the Bilat- eral Multi-perspective Matching framework. Data is collected by mapping the SNLI dataset with the image dataset Flickr30k. The result dataset, made pub- licly available, has more than 565k instances. Experiments on this dataset show that the multimodal model outperforms the state-of-the-art textual model. References 1

http://www.nusl.cz/ntk/nusl-387831

Identifer	oai:union.ndltd.org:nusl.cz/oai:invenio.nusl.cz:387831
Date	January 2018
Creators	Vu Trong, Hoa
Contributors	Pecina, Pavel, Libovický, Jindřich
Source Sets	Czech ETDs
Language	English
Detected Language	English
Type	info:eu-repo/semantics/masterThesis
Rights	info:eu-repo/semantics/restrictedAccess

Page generated in 0.002 seconds

Vyvozování v přirozeném jazyce s využitím obrazových dat / Grounding Natural Language Inference on Images

Description

Links & Downloads

Tags

Additional Fields