Explainable Network Traffic Classification Using XGBoost and SHAP: Interpreting Feature Contributions Across Traffic Classes

Authors

DOI:

https://doi.org/10.26636/jtit.2026.3.2724

Keywords:

deep learning, explainable AI, machine learning, network traffic classification

Abstract

Network traffic classification has become an essential component of modern network management, cyber security, and quality-of-service provisioning. The widespread adoption of encryption technologies has reduced the effectiveness of traditional traffic identification techniques, leading to the use of ML approaches based on flow-level characteristics. Although several machine learning and deep learning models were evaluated, the black-box nature complicates the practical deployment in security critical environments. This research proposes a framework for ML-based network traffic classification of explainable artificial intelligence (XAI). The CIC-Darknet2020 dataset is divided into four classes: non-TOR, non-VPN, TOR and VPN. Additionally, machine learning algorithms, including ID3, k-nearest neighbors (KNN), random forest, CatBoost, XGBoost, and LSTM are used as a deep learning approach. The evaluation is carried out through stratified split of trains and 10-fold cross-validation, while XGBoost is classified for explainability analysis. By ensuring model transparency, SHAP (SHapley Additive exPlanations) identifies the most influential features contributing to classification predictions. Furthermore, a novel category SHAP analysis is introduced by grouping higher-level behavioral categories, including temporal, statistical, rate-based, and TCP-related features. The results revealed that temporal traffic characteristics and transport layer behavioral features influence classification outcomes, particularly for encrypted traffic classes. Moreover, the framework demonstrates that high classification performance and model interpretability are achievable simultaneously within a category-level study that enhances the transparency, trustworthiness, and practical applicability of machine learning-based network traffic classification systems.

Downloads

Download data is not yet available.

References

[1] D. Sarkar, P. Vinod and S.Y. Yerima, "Detection of Tor Traffic using Deep Learning", 2020 IEEE/ACS 17th International Conference on Computer Systems and Applications (AICCSA), Antalya, Turkey, 2020. DOI: https://doi.org/10.1109/AICCSA50499.2020.9316533
View in Google Scholar

[2] R. Venkateswaran, "Virtual Private Networks", IEEE Potentials, vol. 20, pp. 11-15, 2001. DOI: https://doi.org/10.1109/45.913204
View in Google Scholar

[3] Cloudflare, "Worldwide Overview: Cloudflare Radar". [Online] Available: https://radar.cloudflare.com [Accessed: 10.07.2026].
View in Google Scholar

[4] V. Gautam and S. Ganguli, Securing Digital India Against AI and Cyber Threats, Zscaler, 2026 (ISBN: 9798903339068).
View in Google Scholar

[5] Y. Chandiramani, "Attention aux Fausses Versions de Super Mario Run sur Android!", Global Security Mag Online, 2017 https://www.globalsecuritymag.fr/Yogi-Chandiramani-Zscaler,20170116,68299.html (in French).
View in Google Scholar

[6] S. Mali, M. Gujral, and A.K. Cherukuri, "Encrypted Network Traffic Classification Using Intelligent Techniques", Cureus Journal of Computer Science, 2025. DOI: https://doi.org/10.7759/s44389-024-02701-2
View in Google Scholar

[7] G. Aceto et al., "Characterization and Prediction of Mobile-app Traffic Using Markov Modeling", IEEE Transactions on Network and Service Management, vol. 18, pp. 907-925, 2021. DOI: https://doi.org/10.1109/TNSM.2021.3051381
View in Google Scholar

[8] M. Lotfollahi, R.S.H. Zade, M.J. Siavoshani, and M. Saberian, "Deep Packet: A Novel Approach for Encrypted Traffic Classification Using Deep Learning", ArXiv, 2018. DOI: https://doi.org/10.1007/s00500-019-04030-2
View in Google Scholar

[9] M. Ghaleb et al., "Explainable AI for Lightweight Network Traffic Classification Using Depthwise Separable Convolutions", IEEE Open Journal of the Computer Society, vol. 6, pp. 908-920, 2025,. DOI: https://doi.org/10.1109/OJCS.2025.3576495
View in Google Scholar

[10] H. Hagras, "Towards Human Understandable Explainable AI", Computer, vol. 51, pp. 28-36, 2018. DOI: https://doi.org/10.1109/MC.2018.3620965
View in Google Scholar

[11] S.M. Lundberg and S.-I. Lee, "A Unified Approach to Interpreting Model Predictions", ArXiv, 2017.
View in Google Scholar

[12] M.T. Ribeiro, S. Singh, and C. Guestrin, "Why Should I Trust You?: Explaining the Predictions of Any Classifier", Proc. of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 1135-1144, 2016. DOI: https://doi.org/10.1145/2939672.2939778
View in Google Scholar

[13] M. Sundararajan, A. Taly, and Q. Yan, "Axiomatic Attribution for Deep Networks", ArXiv, 2017.
View in Google Scholar

[14] A. Shrikumar, P. Greenside, and A. Kundaje, "Learning Important Features Through Propagating Activation Differences", ArXiv, 2019.
View in Google Scholar

[15] M.D. Zeiler and R. Fergus, "Visualizing and Understanding Convolutional Networks", Computer Vision - ECCV 2014, pp. 818-833, 2014. DOI: https://doi.org/10.1007/978-3-319-10590-1_53
View in Google Scholar

[16] A. Nascita et al., "XAI Meets Mobile Traffic Classification: Understanding and Improving Multimodal Deep Learning Architectures", IEEE Transactions on Network and Service Management, vol. 18, pp. 4225-4246, 2021. DOI: https://doi.org/10.1109/TNSM.2021.3098157
View in Google Scholar

[17] K. Singh, A. Kashyap, and A.K. Cherukuri, "Interpretable Anomaly Detection in Encrypted Traffic Using SHAP with Machine Learning Models", ArXiv, 2025.
View in Google Scholar

[18] S.N. Zeleke, A.F. Jember, and M. Bochicchio, "Integrating Explainable AI for Effective Malware Detection in Encrypted Network Traffic", ArXiv, 2025.
View in Google Scholar

[19] A.N. Gummadi, J.C. Napier, and M. Abdallah, "XAI-IoT: An Explainable AI Framework for Enhancing Anomaly Detection in IoT Systems", IEEE Access, vol. 12, pp. 71024-71054, 2024. DOI: https://doi.org/10.1109/ACCESS.2024.3402446
View in Google Scholar

[20] K. Alam et al., "SXAD: Shapely eXplainable AI-Based Anomaly Detection Using Log Data", IEEE Access, vol. 12, pp. 95659-95672, 2024. DOI: https://doi.org/10.1109/ACCESS.2024.3425472
View in Google Scholar

[21] Canadian Institute for Cybersecurity (CIC), "CIC-Darknet2020", University of New Brunswick. [Online] Available: https://www.unb.ca/cic/datasets/darknet2020.html [Accessed: 10.07.2026].
View in Google Scholar

[22] Canadian Institute for Cybersecurity (CIC), "Tor-nonTor dataset (ISCXTor2016)", University of New Brunswick. [Online] Available: https://www.unb.ca/cic/datasets/tor.html [Accessed: 10.07.2026].
View in Google Scholar

[23] Canadian Institute for Cybersecurity (CIC), "VPN-nonVPN dataset (ISCXVPN2016)", University of New Brunswick. [Online] Available: https://www.unb.ca/cic/datasets/vpn.html [Accessed: 10.07.2026].
View in Google Scholar

[24] J.R. Almonteros and J.B. Matias, "Integration of Stratified KFold Cross Validation to Enhance Prediction Accuracy: A Comparison Study", 2024 5th International Conference on Data Analytics for Business and Industry (ICDABI), Zallaq, Bahrain, 2024. DOI: https://doi.org/10.1109/ICDABI63787.2024.10800425
View in Google Scholar

[25] O. Oyedele, "Determining the Optimal Number of Folds to Use in a K-fold Cross-validation: A Neural Network Classification Experiment", Research in Mathematics, vol. 10, art. no. 2201015, 2023. DOI: https://doi.org/10.1080/27684830.2023.2201015
View in Google Scholar

Downloads

Submitted

2026-07-13

Published

2026-09-30

Issue

Section

ARTICLES FROM THIS ISSUE

How to Cite

[1]
B. S. Deghem, “Explainable Network Traffic Classification Using XGBoost and SHAP: Interpreting Feature Contributions Across Traffic Classes”, JTIT, vol. 105, no. 3, pp. 114–122, Sep. 2026, doi: 10.26636/jtit.2026.3.2724.