MULTILINGUAL CYBERBULLYING DETECTION SYSTEM

Pawar, Rohit Sidram

doi:10.25394/PGS.8035463.v1

MULTILINGUAL CYBERBULLYING DETECTION SYSTEM.pdf (1.13 MB)

MULTILINGUAL CYBERBULLYING DETECTION SYSTEM

thesis

posted on 2019-06-11, 14:50 authored by Rohit Sidram PawarRohit Sidram Pawar

Since the use of social media has evolved, the ability of its users to bully others has increased. One of the prevalent forms of bullying is Cyberbullying, which occurs on the social media sites such as Facebook©, WhatsApp©, and Twitter©. The past decade has witnessed a growth in cyberbullying – is a form of bullying that occurs virtually by the use of electronic devices, such as messaging, e-mail, online gaming, social media, or through images or mails sent to a mobile. This bullying is not only limited to English language and occurs in other languages. Hence, it is of the utmost importance to detect cyberbullying in multiple languages. Since current approaches to identify cyberbullying are mostly focused on English language texts, this thesis proposes a new approach (called Multilingual Cyberbullying Detection System) for the detection of cyberbullying in multiple languages (English, Hindi, and Marathi). It uses two techniques, namely, Machine Learning-based and Lexicon-based, to classify the input data as bullying or non-bullying. The aim of this research is to not only detect cyberbullying but also provide a distributed infrastructure to detect bullying. We have developed multiple prototypes (standalone, collaborative, and cloud-based) and carried out experiments with them to detect cyberbullying on different datasets from multiple languages. The outcomes of our experiments show that the machine-learning model outperforms the lexicon-based model in all the languages. In addition, the results of our experiments show that collaboration techniques can help to improve the accuracy of a poor-performing node in the system. Finally, we show that the cloud-based configurations performed better than the local configurations.

History

Degree Type

Master of Science

Department

Computer Science

Campus location

Indianapolis

Advisor/Supervisor/Committee Chair

Dr. Rajeev R. Raje

Additional Committee Member 2

Dr. Mihran Tuceryan

Additional Committee Member 3

Dr. Arjan Durresi

Usage metrics

Keywords

Distributed computing Natural language processsing machine Learning Predictions cloud applications Indian languages Computer Software Computer Engineering Computer System Architecture Distributed Computing Knowledge Representation and Machine Learning Natural Language Processing

Licence

CC BY 4.0

Exports

RefWorks

BibTeX

Ref. manager

Endnote

DataCite

NLM

DC

MULTILINGUAL CYBERBULLYING DETECTION SYSTEM

History

Degree Type

Department

Campus location

Advisor/Supervisor/Committee Chair

Additional Committee Member 2

Additional Committee Member 3

Usage metrics

Categories

Keywords

Licence

Exports