Enumerating Trillion Subgraphs On Distributed Systems

Overview

How can we find patterns from an enormous graph with billions of vertices and edges? The subgraph enumeration, which is to find patterns from a graph, is an important task for graph data analysis with many applications, including analyzing the social network evolution, measuring the significance of motifs in biological networks, observing the dynamics of Internet, and so on. Especially, the triangle enumeration, a special case of the subgraph enumeration, where the pattern is a triangle, has many applications such as identifying suspicious users in social networks, detecting web spams, and finding communities. However, recent networks are so large that most of the previous algorithms fail to process them. Recently, several MapReduce algorithms have been proposed to address such large networks; however, they suffer from the massive shuffled data resulting in a very long processing time.

In this article, we propose scalable methods for enumerating trillion subgraphs on distributed systems. We first propose PTE (Pre-partitioned Triangle Enumeration), a new distributed algorithm for enumerating triangles in enormous graphs by resolving the structural inefficiency of the previous MapReduce algorithms. PTE enumerates trillions of triangles in a billion scale graph by decreasing three factors: the amount of shuffled data, total work, and network read. We also propose PSE (Pre-partitioned Subgraph Enumeration), a generalized version of PTE for enumerating subgraphs that match an arbitrary query graph. Experimental results show that PTE provides 79 times faster performance than recent distributed algorithms on real-world graphs, and succeeds in enumerating more than 3 trillion triangles on the ClueWeb12 graph with 6.3 billion vertices and 72 billion edges. Furthermore, PSE successfully enumerates 265 trillion clique subgraphs with 4 vertices from a subdomain hyperlink network, showing 47 times faster performance than the state of the art distributed subgraph enumeration algorithm.

Paper

Our proposed method is described in the following paper.

Enumerating Trillion Subgraphs On Distributed Systems
Ha-Myung Park, Francesco Silvestri, Rasmus Pagh, Chin-Wan Chung, Sung-Hyon Myaeng, U Kang.
TKDD '18.
[PDF] [BIBTEX]

Code

PSE for Hadoop: pse-1.0.tar.gz [download]
PSE for Spark: pse-spark-1.0.tar.gz [download]

Datasets

Name	#Nodes	#Edges	Description	Source
Skitter	1.7M	11M	Internet topology graph	SNAP - Skitter
Youtube	3.2M	12M	Friendship network in YouTube	Konect
LiveJournal	4.8M	69M	LiveJournal online social network	SNAP - LiveJornal1
Orkut	3.1M	117M	Orkut online social network	SNAP - Orkut
Twitter	42M	1.2B	Following network of Twitter	Kwak10www - Twitter
Friendster	66M	1.8B	Friendster online social network	SNAP - Friendster
SubDomain	0.1B	1.9B	Links among subdomains on the Web	WDC - Subdomain/Host Graph
YahooWeb	1.4B	6.6B	Page level hyperlink network on the Web	Yahoo-webscope
ClueWeb09	4.8B	7.9B	Page level hyperlink network on the Web	ClueWeb09 Wiki
ClueWeb12	6.3B	72B	Page level hyperlink network on the Web	ClueWeb12 Web Graph

People

Ha-Myung Park (Korea Advanced Institute of Science and Technology, Korea)
Francesco Silvestri (University of Padova)
Rasmus Pagh (IT University of Copenhagen)
Chin-Wan Chung (Chongqing University of Technology and KAIST)
Sung-Hyon Myaeng (Korea Advanced Institute of Science and Technology, Korea)
U Kang (Seoul National University, Korea)