Shrcse
Online Gambling Website Detection Using Gated Recurrent Unit
Data Science
NLP
Web Crawling
Streamlit
Scrapy
"A web-based solution designed to detect and classify online gambling sites, utilizing techniques such as NLP and web crawling for effective detection."
Introduction
"The internet in Indonesia has revolutionized daily life, boosting fields like the economy, society, and culture. But like any powerful tool, it's got a dark side—cue online gambling. Imagine games and lotteries where people bet money or valuables, all just a click away. Sounds thrilling? Not when it leads to addiction, lost savings, data breaches, and weakened social values. It's also illegal, violating Article 27(2) of Law No. 19 of 2016. The government's been busy, blocking over 60,000 pieces of online gambling content in 2023 alone, but the problem keeps evolving. It's like a game of whack-a-mole, and it's clear we need smarter, more effective solutions to tackle this ever-changing issue."
Methodology
Thisprojectaimstodevelopaweb-basedsystemfordetectingonlinegamblingwebsites,utilizingScrapy,Streamlit,andGRUmodels.
The dataset was created by Crawling 400 URLs, with the collected data manually labeled to prepare it for subsequent analysis.
1
1
To help the models grasp the context of online gambling, a Word2vec model was developed to generate meaningful word embeddings.
2
2
Using the prepared data and word embeddings, a GRU and BiGRU model was trained with defined parameters.
3
3
Selected GRU model was utilized to drive the web-based detection system, integrating all previous steps into a unified implementation.
4
4
The implementation achieved peak performance, delivering a 100% detection success rate when tested on a dataset of 200 URLs.
5
5
Result
The following demonstrates the web-based implementation designed to streamline the entire process in a single operation. This includes inputting the URL for detection, crawling the website, extracting and preprocessing the text, running detection using GRU models, and making the final decision.
How It Work
UponenteringthetargetwebsiteURL,abotspideragentisdeployedtonavigatethesite,systematicallycrawlingandextractingalltextualcontent.Thecollecteddataiscompiledintoasingledataset.However,thisrawdataistypicallyunrefinedandnecessitatespreprocessingsteps,includingtextlengthrestriction,filtering,casenormalization,andtheremovalofstopwords,toenhanceitsqualityforsubsequentanalysis.
EachlineoftextinthedatasetissubsequentlyclassifiedusingthedevelopedGRUmodel,whichassignsascorerangingfrom0to1basedonthecontent'scontext.Textcontaininggambling-relatedcontentreceivesahigherscore,approaching1,whereasnon-gamblingcontentisassignedalowerscore,closerto0.
Afterclassifyingallthetext,adecisionismadebycalculatingthetotalscoreacrossthedataset.Severalthresholdvalues—25%,50%,75%,and90%—weretestedtodeterminethemosteffectivecutofffordistinguishingbetweengamblingandnon-gamblingwebsites.The50%thresholdemergedastheoptimalfit,achieving0%falsepositivesand0%falsenegatives.Ifthetotalscoreofawebsiteexceeds50%,itisclassifiedasagamblingwebsite;otherwise,itisdeemednon-gambling.Theresultsarethenpresented,andtheURLisstoredinthedatabase.
Snapshot



