International Journal of Advances in Data and Information Systems
Vol. 2 No. 1 (2021): April 2021 - International Journal of Advances in Data and Information Systems

Features Selection for Entity Resolution in Prostitution on Twitter

Reisa Permatasari (Departement of Information System, Institut Teknologi Sepuluh Nopember)
Nur Aini Rakhmawati (Departement of Information System, Institut Teknologi Sepuluh Nopember)



Article Info

Publish Date
27 Mar 2021

Abstract

Entity resolution is the process of determining whether two references to real-world objects refer to the same or different purposes. This study applies entity resolution on Twitter prostitution dataset based on features with the Regularized Logistic Regression training and determination of Active Learning on Dedupe and based on graphs using Neo4j and Node2Vec. This study found that maximum similarity is 1 when the number of features (personal, location and bio specifications) is complete. The minimum similarity is 0.025662627 when the amount of harmful training data. The most influencing similarity feature is the cellphone number with the lowest starting range from 0.997678459 to 0.999993523.  The parameter - length of walk per source has the effect of achieving the best similarity accuracy reaching 71.4% (prediction 14 and yield 10).

Copyrights © 2021






Journal Info

Abbrev

IJADIS

Publisher

Subject

Computer Science & IT Electrical & Electronics Engineering

Description

International Journal of Advances in Data and Information Systems (IJADIS) (e-ISSN: 2721-3056) is a peer-reviewed journal in the field of data science and information system that is published twice a year; scheduled in April and October. The journal is published for those who wish to share ...