Home  /  Entropy  /  Vol: 19 Núm: 12 Par: Decembe (2017)  /  Article

Do We Really Need to Catch Them All? A New User-Guided Social Media Crawling Method


[-15]With the growing use of popular social media services like Facebook and Twitter it is challenging to collect all content from the networks without access to the core infrastructure or paying for it. Thus, if all content cannot be collected one must consider which data are of most importance. In this work we present a novel User-guided Social Media Crawling method (USMC) that is able to collect data from social media, utilizing the wisdom of the crowd to decide the order in which user generated content should be collected to cover as many user interactions as possible. USMC is validated by crawling 160 public Facebook pages, containing content from 368 million users including 1.3 billion interactions, and it is compared with two other crawling methods. The results show that it is possible to cover approximately 75% of the interactions on a Facebook page by sampling just 20% of its posts, and at the same time reduce the crawling time by 53%. In addition, the social network constructed from the 20% sample contains more than 75% of the users and edges compared to the social network created from all posts, and it has similar degree distribution.

 Articles related

Opik Taupik Kurahman,Tedi Priatna,Tri Cahyanto    

Law of the Republic of Indonesia Number 33 Year 2014 on Halal Product Warranty (JPH)  mandates the Government that the product circulating in Indonesia is guaranteed halal by the Halal Product Security Organizing Body (BPJPH). Public participat... see more

Muhammad Nurul Huda,Muhammad Burhan,Ahmad Satibi,Helmy Agus Pradita,Aries Saifudin,Irpan Kusyadi    

Testing an application serves to check whether an application / program is running with what is desired or whether there are still errors / errors that need to be corrected so that the program created into a quality program. Software testing method consi... see more

Yulianti Yulianti,Anif Biantoro,Geri Santoso Adi,Muhammad Asshidiqie,Rizky Abiansyah Putra,Aries Saifudin    

This research describes implementation of information technology in  elementary school library. library strategy utilizes an electronic technique that isn't yet incorporated, and the information handling utilizes MS Office to store understudy inform... see more

Suniwati .,Indri Astuti,Afandi .    

Abstrak: Pada masa sekarang ini siswa dengan kebutuhan khusus atau yang biasa juga di sebut disabilitas telah banyak bisa kita jumpai. Kelompok ini terdiri dari kebutuhan khusus pada fisik dan psikis mereka. Anak-anak dengan kebutuhan khusus ini memang s... see more

Amri Haq Anugrah,Gunarhadi Gunarhadi,Tri Rejeki Andayani    

Learning media is one of the important things in the child's learning process. The purpose of this study was to determine the teacher's needs for learning media for children with special needs of the deaf type in learning hijaiyah letters. This research ... see more