Discovering maximal motif cliques in large heterogeneous information networks
Accepted version
Peer-reviewed
Repository URI
Repository DOI
Change log
Authors
Abstract
We study the discovery of cliques (or "complete" subgraphs) in heterogeneous information networks (HINs). Existing clique-finding solutions often ignore the rich semantics of HINs. We propose motif clique, or m-clique, which redefines subgraphs completeness with respect to a given motif. A motif essentially a small subgraph pattern, is a fundamental building block of an HIN. The m-clique concept is general and allows us to analyse "complete" subgraphs in an HIN with respect to desired high-order connection patterns. We further investigate the maximal m-clique enumeration problem (MMCE), which finds all maximal m-cliques not contained in any other m-cliques. Because MMCE is NP-hard, developing an accurate and efficient solution for MMCE is not straightforward. we thus present the META algorithm, which employs advanced pruning strategies to effectively reduce the search space. We also design fast techniques to avoid generating duplicated maximal m-clique instances. Our extensive experiments on large real and synthetic HINs how that META is highly effective and efficient.
Description
Keywords
Journal Title
Conference Name
Journal ISSN
Volume Title
Publisher
Publisher DOI
Sponsorship
Wellcome Trust (100574/Z/12/Z)
Medical Research Council (MC_UU_12012/1)
Medical Research Council (MC_UU_12012/5)
Medical Research Council (MC_PC_12012)