A Hybrid BERTopic–NAMAA Framework for Arabic News Topic Modeling, Classification, and Evaluation

Authors

  • Noor S. Dawood
  • Salma A. Mahmood

DOI:

https://doi.org/10.22399/ijnasen.38

Keywords:

Arabic Natural Language Processing, BERTopic, Topic Modeling; Arabic, Text Classification, Arabic News, AraBERT, NAMAA, Cohen's Kappa

Abstract

Topic detection is a powerful technique for uncovering the latent semantic structures in large-scale unlabeled text. Yet, the application to Arabic is still an open challenge due to the complex morphology of the language and the absence of comprehensive solution that extends topic discovery by systematic evaluation. In this paper, a novel hybrid architecture based on BERTopic (for unsupervised topic detection) in conjunction with NAMAA (an Arabic classification technique) for topic category assignment and external evaluation is presented. The suggested architecture involves an Arabic-specific preprocessing layer, a contextual document embedding through four Arabic pre-trained transformer models (AraBERTv0.2, AraBERTv2, Asafaya, and QARiB), a topic extraction layer by utilizing the BERTopic model, an automatic topic category assignment layer via NAMAA, and an evaluation step for the entire architecture based on both intrinsic (NPMI, C_v coherence, and Topic Diversity) and extrinsic (Accuracy and Cohen's Kappa) metrics.          The proposed framework was tested on two Arabic news related Economy and Sports domains: an imbalanced dataset (D1) of 7,212 documents and a balanced dataset (D2) of 3,158 documents. BERTopic yielded between 59 and 116 topics on D1 and between 62 and 125 topics on D2. On the balanced dataset Asafaya scored the best C_v coherence (0.579) and Cohen Kappa (0.937), while AraBERTv2 ranked the best NPMI (0.289) and Topic Diversity (0.702). The best classification accuracy with Asafaya on D2 was 96.8%, while 98.2 with AraBERTv2 on D1. To sum up, these results indicate that our framework delivers a robust and efficient solution for Arabic news topic discovery and evaluation, revealing also the impact of embedding model choice and class distribution on topic quality, classification accuracy and agreement with the original dataset labels.

References

[1] A. M. Alayba, "Arabic Natural Language Processing (NLP): A Comprehensive Review of Challenges, Techniques, and Emerging Trends," Computers, vol. 14, no. 11, p. 497, Nov. 2025, doi: 10.3390/computers14110497.

[2] S. V. Mahadevkar, S. Patil, K. Kotecha, L. W. Soong, and T. Choudhury, "Exploring AI-driven approaches for unstructured document analysis and future horizons," J Big Data, vol. 11, no. 1, p. 92, Jul. 2024, doi: 10.1186/s40537-024-00948-z.

[3] C. D. Manning, P. Raghavan, and H. Schütze, An Introduction to Information Retrieval. Cambridge, England: Cambridge University Press, 2008.

[4] G. Salton and C. Buckley, "Term Weighting Approaches in Automatic Text Retrieval," Information Processing & Management, vol. 24, no. 5, pp. 513–523, 1988.

[5] D. M. Blei, A. Y. Ng, and M. I. Jordan, "Latent Dirichlet Allocation," Journal of Machine Learning Research, 2003. [Online]. Available: https://www.jmlr.org/papers/v3/blei03a.html

[6]A. Vaswani et al., "Attention Is All You Need," Aug. 02, 2023, arXiv:1706.03762. doi: 10.48550/arXiv.1706.03762.

[7] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding," doi: 10.48550/arXiv.1810.04805.

[8] N. Habash, Arabic Natural Language Processing. United States: Morgan and Claypool, 2009.

[9]D. M. Blei, "Probabilistic topic models," Commun. ACM, vol. 55, no. 4, pp. 77–84, Apr. 2012, doi: 10.1145/2133806.2133826.

[10] V. Keselj, "Book Review: Speech and Language Processing (second edition) by Daniel Jurafsky and James H. Martin," Computational Linguistics, vol. 35, no. 3, Sep. 2009, doi: 10.1162/coli.B09-001.

[11] M. Grootendorst, "BERTopic: Neural topic modeling with a class-based TF-IDF procedure."

[12] O. Elshehy, O. Nacar, A. Djamai, M. Ragab, K. A. Jallad, and M. Abdelazim, "AraModernBERT: Transtokenized Initialization and Long-Context Encoder Modeling for Arabic," 2026, arXiv. doi: 10.48550/ARXIV.2603.09982.

[13] A. Abuzayed and H. Al-Khalifa, "BERT for Arabic Topic Modeling: An Experimental Study on BERTopic Technique," Procedia Computer Science, vol. 189, pp. 191–194, 2021, doi: 10.1016/j.procs.2021.05.096.

[14] A. Abdelrazek, W. Medhat, E. Gawish, and A. Hassan, "Topic Modeling on Arabic Language Dataset: Comparative Study," in Advances in Model and Data Engineering in the Digitalization Era, vol. 1751, P. Fournier-Viger, A. Hassan, L. Bellatreche, A. Awad, A. Ait Wakrime, Y. Ouhammou, and I. Ait Sadoune, Eds., Communications in Computer and Information Science, vol. 1751. Cham: Springer Nature Switzerland, 2022, pp. 61–71. doi: 10.1007/978-3-031-23119-3_5.

[15] S. Aouichaty, Y. Maleh, M. T. Mohtadi, A. Hajami, and H. Allali, "Sustainable Topic Modeling for Legal Moroccan Arabic Language: A Challenging Study on BERTopic Technique," Procedia Computer Science, vol. 236, pp. 582–588, 2024, doi: 10.1016/j.procs.2024.05.069.

[16] S. Ben Ali, Z. Kechaou, and A. Wali, "Arabic fake news detection in social media Based on AraBERT," in 2022 IEEE 21st International Conference on Cognitive Informatics & Cognitive Computing (ICCI*CC), Toronto, ON, Canada: IEEE, Dec. 2022, pp. 214–220. doi: 10.1109/ICCICC57084.2022.10101635.

[17] A. B. Nassif, A. Elnagar, O. Elgendy, and Y. Afadar, "Arabic fake news detection based on deep contextualized embedding models," Neural Comput & Applic, vol. 34, no. 18, pp. 16019–16032, Sep. 2022, doi: 10.1007/s00521-022-07206-4.

[18] M. Abdul-Mageed, A. Elmadany, and E. M. B. Nagoudi, "ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic," in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 2021, pp. 7088–7105. doi: 10.18653/v1/2021.acl-long.551.

[19] O. Karajeh, M. N. Al-Kabi, and E. A. Fox, "Fusing AraBERT and Graph Neural Networks for Enhanced Arabic Text Classification," in 2023 24th International Arab Conference on Information Technology (ACIT), Ajman, United Arab Emirates: IEEE, Dec. 2023, pp. 1–8. doi: 10.1109/ACIT58888.2023.10453909.

[20] M. Z. Al-Taie, "Comparative Study of Machine Learning Approaches for Detecting Fake News in Arabic Text," IETI Trans. Data Anal. Forecast.vol. 3, no. 1, pp. 18–31, May 2025, doi: 10.3991/itdaf.v3i1.53575.

[21] M. Abbas and K. Smaili, "Comparison of Topic Identification Methods for Arabic Language," presented at the RANLP 2005: Recent Advances in Natural Language Processing, 2005, pp. 14–17.

[22] O. Einea, A. Elnagar, and R. Al Debsi, "SANAD: Single-label Arabic News Articles Dataset for automatic text categorization," Data in Brief, vol. 25, p. 104076, Aug. 2019, doi: 10.1016/j.dib.2019.104076.

[23] L. McInnes, J. Healy, and J. Melville, "UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction."

[24] L. McInnes, J. Healy, and S. Astels, "hdbscan: Hierarchical density based clustering," JOSS, vol. 2, no. 11, p. 205, Mar. 2017, doi: 10.21105/joss.00205.

[25] W. Antoun, F. Baly, and H. Hajj, "AraBERT: Transformer-based Model for Arabic Language Understanding."

[26] A. Safaya, M. Abdullatif, and D. Yuret, "KUISAIL at SemEval-2020 Task 12: BERT-CNN for Offensive Speech Identification in Social Media," in Proceedings of the Fourteenth Workshop on Semantic Evaluation, Barcelona (online): International Committee for Computational Linguistics, 2020, pp. 2054–2059. doi: 10.18653/v1/2020.semeval-1.271.

[27] A. Abdelali, S. Hassan, H. Mubarak, K. Darwish, and Y. Samih, "Pre-Training BERT on Arabic Tweets: Practical Considerations."

[28] G. Bouma, "Normalized (Pointwise) Mutual Information in Collocation Extraction."

[29] J. H. Lau, D. Newman, and T. Baldwin, "Machine Reading Tea Leaves: Automatically Evaluating Topic Coherence and Topic Model Quality," in Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, Gothenburg, Sweden: Association for Computational Linguistics, 2014, pp. 530–539. doi: 10.3115/v1/E14-1056.

[30] D. Newman, J. H. Lau, K. Grieser, and T. Baldwin, "Automatic Evaluation of Topic Coherence."

[31] M. Röder, A. Both, and A. Hinneburg, "Exploring the Space of Topic Coherence Measures," in Proceedings of the Eighth ACM International Conference on Web Search and Data Mining, Shanghai, China: ACM, Feb. 2015, pp. 399–408. doi: 10.1145/2684822.2685324.

[32] A. B. Dieng, F. J. R. Ruiz, and D. M. Blei, "Topic Modeling in Embedding Spaces," Transactions of the Association for Computational Linguistics, vol. 8, pp. 439–453, Dec. 2020, doi: 10.1162/tacl_a_00325.

[33] M. Sokolova and G. Lapalme, "A systematic analysis of performance measures for classification tasks," Information Processing & Management, vol. 45, no. 4, pp. 427–437, Jul. 2009, doi: 10.1016/j.ipm.2009.03.002.

[34] C. J. Van Rijsbergen, Information Retrieval, 2nd ed. London: Butterworth, 1979.

[35]"Glossary of Terms," Machine Learning, vol. 30, no. 2–3, pp. 271–274, Feb. 1998, doi: 10.1023/A:1017181826899.

[36] J. Cohen, "A Coefficient of Agreement for Nominal Scales," Educational and Psychological Measurement, vol. 20, no. 1, pp. 37–46, 1960.

[37] J. R. Landis and G. G. Koch, "The Measurement of Observer Agreement for Categorical Data," Biometrics, vol. 33, no. 1, p. 159, Mar. 1977, doi: 10.2307/2529310.

Downloads

Published

2026-08-08

How to Cite

Noor S. Dawood, & Salma A. Mahmood. (2026). A Hybrid BERTopic–NAMAA Framework for Arabic News Topic Modeling, Classification, and Evaluation. International Journal of Natural-Applied Sciences and Engineering, 4(1). https://doi.org/10.22399/ijnasen.38

Issue

Section

Articles