Christian Putra's portfolio
← back to all projects
Academic Research•Machine Learning · Deep Learning · Cybersecurity · Research

Phishing Detection: Classical vs Deep Learning

Studi komparatif mengevaluasi Random Forest dan hybrid CNN-LSTM untuk deteksi URL phishing menggunakan dataset LegitPhish.

PythonScikit-LearnTensorFlowSMOTEKeras
Dataset Size
101,219 URLs
RF Accuracy
99.98%
CNN-LSTM Accuracy
99.97%
ROC-AUC
1.00

overview

Phishing attacks remain a critical cybersecurity threat, while traditional blocklist-based methods struggle against rapidly evolving malicious URLs[cite: 6]. This study presents a controlled comparative analysis between a traditional Machine Learning model (Random Forest) and a Deep Learning model (hybrid CNN-LSTM) under identical experimental conditions[cite: 6].

the problem

Modern phishing campaigns deploy complex URLs featuring excessive subdomains, deep directory paths, and randomized strings designed to bypass simple filters[cite: 6]. Security systems require accurate automated classifiers, yet balancing detection performance, computational resource efficiency, and model interpretability remains a major challenge[cite: 6].

dataset & preprocessing

The research utilizes the public LegitPhish dataset from Mendeley Data containing 101,219 labeled URLs (63,678 phishing and 37,540 legitimate) with 16 structural and lexical features[cite: 6]. Data preprocessing involved cleaning missing values, applying MinMaxScaler for numerical lexical attributes, and utilizing the SMOTE technique exclusively on training splits (80:20 ratio) to address class imbalance[cite: 6].

two approaches

The Random Forest model leveraged 16 engineered lexical features optimized via GridSearchCV (n_estimators=100, max_depth=20, min_samples_split=2, criterion=entropy) with 5-fold cross-validation[cite: 6]. In parallel, the hybrid CNN-LSTM deep learning architecture implemented a dual-branch feature fusion mechanism: processing character-level tokenized raw URL sequences through an embedding layer, Conv1D, and LSTM layers, combined with scaled lexical features through a dense layer and concatenation mechanism[cite: 6].

results

Both models achieved exceptional performance with an ROC-AUC of 1.00[cite: 6]. Random Forest slightly outperformed CNN-LSTM in overall accuracy (99.98% vs 99.97%) with only 4 total misclassifications (2 False Positives, 2 False Negatives)[cite: 6]. The hybrid CNN-LSTM achieved 99.97% accuracy, maintaining a very low False Positive Rate (FPR) of under 1% (0.99%)[cite: 6].

interpretability

Feature importance analysis for Random Forest revealed a strong dependency on structural attributes, with path_length (feature 10) acting as the absolute dominant indicator for identifying malicious URLs, followed by domain_name_length and url_entropy[cite: 6]. While Random Forest offers high transparency, CNN-LSTM acts largely as a black-box system but excels in autonomously learning complex sequential patterns[cite: 6].

limitations

The study acknowledges several limitations: evaluation was restricted to a single split of the LegitPhish dataset without external or temporal validation, no formal statistical significance tests (such as McNemar's test) were performed, and computational resource benchmarking (inference latency/memory) was not measured[cite: 6].