Show simple item record

dc.contributor.authorLitschko, Robert
dc.contributor.authorGlavas, Goran
dc.contributor.authorPonzetto, Simone Paolo
dc.contributor.authorVulic, Ivan
dc.date.accessioned2018-09-05T12:42:17Z
dc.date.available2018-09-05T12:42:17Z
dc.date.issued2018
dc.identifier.urihttps://www.repository.cam.ac.uk/handle/1810/279400
dc.description.abstractWe propose a fully unsupervised framework for ad-hoc cross-lingual information retrieval (CLIR) which requires no bilingual data at all. The framework leverages shared cross-lingual word embedding spaces in which terms, queries, and documents can be represented, irrespective of their actual language. The shared embedding spaces are induced solely on the basis of monolingual corpora in two languages through an iterative process based on adversarial neural networks. Our experiments on the standard CLEF CLIR collections for three language pairs of varying degrees of language similarity (English-Dutch/Italian/Finnish) demonstrate the usefulness of the proposed fully unsupervised approach. Our CLIR models with unsupervised cross-lingual embeddings outperform baselines that utilize cross-lingual embeddings induced relying on word-level and document-level alignments. We then demonstrate that further improvements can be achieved by unsupervised ensemble CLIR models. We believe that the proposed framework is the first step towards development of effective CLIR models for language pairs and domains where parallel data are scarce or non-existent.
dc.publisherACM
dc.subjectUnsupervised cross-lingual IR
dc.subjectcross-lingual vector spaces
dc.titleUnsupervised Cross-Lingual Information Retrieval Using Monolingual Data Only
dc.typeConference Object
prism.endingPage1256
prism.publicationDate2018
prism.publicationNameACM/SIGIR PROCEEDINGS 2018
prism.startingPage1253
dc.identifier.doi10.17863/CAM.26775
dcterms.dateAccepted2018-04-11
rioxxterms.versionofrecord10.1145/3209978.3210157
rioxxterms.licenseref.urihttp://www.rioxx.net/licenses/all-rights-reserved
rioxxterms.licenseref.startdate2018
rioxxterms.typeConference Paper/Proceeding/Abstract
pubs.funder-project-idEuropean Research Council (648909)
cam.issuedOnline2018-06-27
pubs.conference-nameSIGIR '18: The 41st International ACM SIGIR conference on research and development in Information Retrieval
pubs.conference-start-date2018-07-08
pubs.conference-finish-date2018-05-12
rioxxterms.freetoread.startdate2019-07-08


Files in this item

Thumbnail

This item appears in the following Collection(s)

Show simple item record