<?xml version="1.0" encoding="UTF-8"?>
<doi_batch version="5.4.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns="http://www.crossref.org/schema/5.4.0" xsi:schemaLocation="http://www.crossref.org/schema/5.4.0 https://www.crossref.org/schemas/crossref5.4.0.xsd" xmlns:jats="http://www.ncbi.nlm.nih.gov/JATS1" xmlns:fr="http://www.crossref.org/fundref.xsd" xmlns:ai="http://www.crossref.org/AccessIndicators.xsd" xmlns:rel="http://www.crossref.org/relations.xsd" xmlns:mml="http://www.w3.org/1998/Math/MathML">
  <head>
    <doi_batch_id>NONE</doi_batch_id>
    <timestamp>20260102091335755</timestamp>
    <depositor>
      <depositor_name>wseas/wseas</depositor_name>
      <email_address>content-registration-form+ja@crossref.org</email_address>
    </depositor>
    <registrant>content-registration-form</registrant>
  </head>
  <body>
    <journal>
      <journal_metadata>
        <full_title>WSEAS TRANSACTIONS ON INFORMATION SCIENCE AND APPLICATIONS</full_title>
        <issn media_type="print">1790-0832</issn>
        <issn media_type="electronic">2224-3402</issn>
      </journal_metadata>
      <journal_article>
        <titles>
          <title>AraBART-based Arabic Lemmatization</title>
        </titles>
        <contributors>
          <person_name sequence="first" contributor_role="author">
            <given_name>Soumia</given_name>
            <surname>Afartass</surname>
            <affiliations>
              <institution>
                <institution_name>Laboratory of Research in Computer Science and Telecommunications (LRIT), Faculty of Sciences Mohammed V University in Rabat, MOROCCO</institution_name>
              </institution>
            </affiliations>
          </person_name>
          <person_name sequence="additional" contributor_role="author">
            <given_name>Fadoua Ataa</given_name>
            <surname>Allah</surname>
            <affiliations>
              <institution>
                <institution_name>Computer Science Studies, Information Systems and Communication Center, The Royal Institute of Amazigh Culture in Rabat, MOROCCO</institution_name>
              </institution>
            </affiliations>
          </person_name>
          <person_name sequence="additional" contributor_role="author">
            <given_name>Khalid</given_name>
            <surname>Minaoui</surname>
            <affiliations>
              <institution>
                <institution_name>Laboratory of Research in Computer Science and Telecommunications (LRIT), Faculty of Sciences Mohammed V University in Rabat, MOROCCO</institution_name>
              </institution>
            </affiliations>
          </person_name>
        </contributors>
        <jats:abstract xml:lang="en">
          <jats:p>Arabic, an inflectional language with a rich morphology and complex syntactic structures, demands robust approaches for effective normalization and lemmatization. While, most existing NLP models focus on English, Modern Standard Arabic remains understudied, particularly in lemmatization tasks. In this paper, we introduce AraBART, the first Arabic model to feature an end-to-end pre-trained encoder-decoder, leveraging the BART architecture. We used the Arabic-PADT UD Treebank and Farasa corpus. Performance was assessed with accuracy, precision, recall, and F1 score. The results show that AraBART surpasses strong baselines, including transformer-based models such as AraT5, mT5, and BERT, achieving a 4.72% improvement in lemmatization accuracy. More importantly, AraBART has achieved an accuracy of 94.71%, approaching the performance of Farasa (97.32%). Furthermore, incorporating parts of the Farasa corpus into our training process showed a clear improvement in accuracy, revealing AraBART's effectiveness for broad applications in Arabic NLP tasks, including text summarization and machine translation.</jats:p>
        </jats:abstract>
        <publication_date media_type="print">
          <month>01</month>
          <day>02</day>
          <year>2026</year>
        </publication_date>
        <publication_date media_type="online">
          <month>01</month>
          <day>02</day>
          <year>2026</year>
        </publication_date>
        <pages>
          <first_page>1</first_page>
        </pages>
        <publisher_item>
          <item_number item_number_type="article_number">1</item_number>
        </publisher_item>
        <ai:program name="AccessIndicators">
          <ai:license_ref>https://creativecommons.org/licenses/by/4.0/deed.en_US</ai:license_ref>
        </ai:program>
        <doi_data>
          <doi>10.37394/23209.2026.23.1</doi>
          <resource>https://wseas.com/journals/isa/2026/a025109-001(2026).pdf</resource>
        </doi_data>
        <citation_list>
          <citation key="ref0">
            <unstructured_citation>S. Afartass, F. A. Allah, K. Minaoui, Transformer-based Arabic Language Lemmatization, 3rd International Conference on Embedded Systems and Artificial Intelligence (ESAI), 2024. http://dx.doi.org/10.1109/ESAI62891.2024.10 913868.</unstructured_citation>
          </citation>
          <citation key="ref1">
            <unstructured_citation>Z. Ghosn, Les sites Internet gouvernementaux au Moyen-Orient, The Arab Advisors Group, 2003, [Online]. https://www.persee.fr/ (Accessed: November 5, 2024).</unstructured_citation>
          </citation>
          <citation key="ref2">
            <unstructured_citation>B. Barraud, Tribune des Droits Humains, Censure de l’internet dans les pays arabes, 2006, [Online]. https://www.academia.edu/94916908/ (Accessed: April 20, 2024).</unstructured_citation>
          </citation>
          <citation key="ref3">
            <unstructured_citation>O. Dakkak, N. Ghneim, A. Shalaby, S. Desouki, R. Sonbol, Arabic Language Resources in HIAST, Conference on Arabic Language Resources and Tools, 2009.</unstructured_citation>
          </citation>
          <citation key="ref4">
            <unstructured_citation>A. Alghani, Lexical extraction and morphological and semantic categorization of the Arabic language dictionary (Extraction lexicale et catégorisation morphologique et sémantique du dictionnaire de la langue arabe), 2020, [Online]. https://memoirepfe.fstusmba.ac.ma/book/1454 (Accessed: November 3, 2024).</unstructured_citation>
          </citation>
          <citation key="ref5">
            <unstructured_citation>I. F. Mohammed, Q. A. Dhayef, Derivation and its Effect on Meaning in English and Arabic: A Contrastive Study, International Journal of Linguistics, 2022. DOI: 10.32996/ijls.2022.2.2.20.</unstructured_citation>
          </citation>
          <citation key="ref6">
            <unstructured_citation>D. Tahar, S. Benharzallah, B. Ali, The Impact of Online Indexing in Improving Arabic Information Retrieval Systems, Informatica, Vol. 42, 2018. https://doi.org/10.31449/inf.v42i4.229.</unstructured_citation>
          </citation>
          <citation key="ref7">
            <unstructured_citation>M. Rizi, S. Maazaoui, A New Deep LearningBased Approach for Cloud Service Classification (Une Nouvelle approche basée deep learning pour la classification de services cloud), 2021, [Online]. https://thesesalgerie.com/ (Accessed: July 24, 2024).</unstructured_citation>
          </citation>
          <citation key="ref8">
            <unstructured_citation>A. El-Sawy, H. El-Bakry, M. Loey, N. Mastorakis, An Intelligent Agent Tutor System for Detecting Arabic Children Handwriting Difficulty Based on Immediate Feedback, WSEAS Transactions on Systems, Vol. 15, 2016, pp. 63-72.</unstructured_citation>
          </citation>
          <citation key="ref9">
            <unstructured_citation>A. H. Alshehri, AraQA-BERT: Towards an Arabic Question Answering System using pre-trainedBERT Models, WSEAS Transactions on Information Science and Applications, vol. 21, 2024, pp. 361-373. https://doi.org/10.37394/23209.2024.21.34.</unstructured_citation>
          </citation>
          <citation key="ref10">
            <unstructured_citation>M. Rizi, S. Maazaoui, Une Nouvelle approche basée deep learning pour la classification de services cloud, Artificial Intelligence and Deep Learning: The Concept and the Applications, Vol. 6, 2011. https://doi.org/10.34874/IMIST.PRSM/reinno va-v6i23.53457.</unstructured_citation>
          </citation>
          <citation key="ref11">
            <unstructured_citation>N. Habash, R. Roth, O. Rambow, R. Eskander, N. Tomeh. Morphological Analysis and Disambiguation for Dialectal Arabic. Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2013, pp. 426–432.</unstructured_citation>
          </citation>
          <citation key="ref12">
            <unstructured_citation>M. Diab, Second Generation AMIRA Tools for Arabic Processing: Fast and Robust Tokenization, POS Tagging, and Base Phrase Chunking, Proceedings of the Second International Conference on Arabic Language Resources and Tools, Cairo, Egypt 2009.</unstructured_citation>
          </citation>
          <citation key="ref13">
            <unstructured_citation>A. Pasha, M. Al-Badrashiny, M. Diab, A. El Kholy, R. Eskander, N. Habash, M. Pooleery, O. Rambow, R. Roth. MADAMIRA: A Fast, Comprehensive Tool for Morphological Analysis and Disambiguation of Arabic. Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14), 2014, pp 1094–1101.</unstructured_citation>
          </citation>
          <citation key="ref14">
            <unstructured_citation>H. Mubarak, Build Fast and Accurate Lemmatization for Arabic, QCRI, Hamad Bin Khalifa University (HBKU), Doha, Qatar, 2018. https://aclanthology.org/L18-1181/.</unstructured_citation>
          </citation>
          <citation key="ref15">
            <unstructured_citation>M. Boudchiche, A. Mazroui, A hybrid approach for Arabic lemmatization, Int. J. Speech Technol, Vol. 22, 2018, pp. 568-573. https://doi.org/10.1007/s10772-018-9528-3.</unstructured_citation>
          </citation>
          <citation key="ref16">
            <unstructured_citation>C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, P. J. Liu, Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, Proceedings of the 37th International Conference on Machine Learning, 2020, pp. 1-20.</unstructured_citation>
          </citation>
          <citation key="ref17">
            <unstructured_citation>E. B. Nagoudi, A. Elmadany, M. AbdulMageed, AraT5: Text-to-Text Transformers for Arabic Language Generation, arXiv, 2021, pp. 1-20. https://doi.org/10.48550/arXiv.2109.12068.</unstructured_citation>
          </citation>
          <citation key="ref18">
            <unstructured_citation>L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, C. Raffel, mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, NAACL 2021, 2021, pp. 1-20.</unstructured_citation>
          </citation>
          <citation key="ref19">
            <unstructured_citation>J. Devlin, M. Chang, K. Lee, K. Toutanova, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, arXiv, Vol. 2019, No. 10, 2019, pp. 1-20. https://doi.org/10.48550/arXiv.1810.04805.</unstructured_citation>
          </citation>
          <citation key="ref20">
            <unstructured_citation>Y. Liu, J. Gu, N. Goyal, X. Li, S. Edunov, M. Ghazvininejad, M. Lewis, L. Zettlemoyer, Multilingual denoising pre-training for neural machine translation, Transactions of the Association for Computational Linguistics, Vol. 8, 2020, pp. 726–742. DOI: 10.1162/tacl_a_00343.</unstructured_citation>
          </citation>
          <citation key="ref21">
            <unstructured_citation>H. Le, ChatGPT: Exploring OpenAI’s Language Model, 2023, [Online]. https://hoanhle.github.io/blog/2023/chatgpt/ (Accessed October 17, 2024).</unstructured_citation>
          </citation>
          <citation key="ref22">
            <unstructured_citation>M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, L. Zettlemoyer. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 7871– 7880. https://doi.org/10.48550/arXiv.1910.13461</unstructured_citation>
          </citation>
          <citation key="ref23">
            <unstructured_citation>A. Özcift, K. Akarsu, F. Yumuk, C. Söylemez, Advancing Natural Language Processing (NLP) Applications of Morphologically Rich Languages with Bidirectional Encoder Representations from Transformers (BERT): An Empirical Case Study for Turkish, Automatic Control and Computer Sciences, 2021. https://doi.org/10.1080/00051144.2021.19221 50.</unstructured_citation>
          </citation>
          <citation key="ref24">
            <unstructured_citation>B. Pengili, D. A. Karras, A detailed comparative analysis towards longer-term traffic loads forecasting with Autoregressive Integrated Moving Average (ARIMA/SARIMA) models to improve transportation services, Financial Engineering, Vol. 3, 2025, pp. 147-171. https://doi.org/10.37394/232032.2025.3.13.</unstructured_citation>
          </citation>
          <citation key="ref25">
            <unstructured_citation>T. Keldenich, Recall, Precision, F1 Score: Simple Metric Explanation, Machine Learning, 2021, [Online]. https://insidemachinelearning.com (Accessed April 25, 2024).</unstructured_citation>
          </citation>
          <citation key="ref26">
            <unstructured_citation>M. Ulcar, M. Robnik-Sikonja, Sequence-tosequence pretraining for a less-resourced Slovenian language, Frontiers in Artificial Intelligence, vol. 3, 2023. doi: 10.3389/frai.2023.932519.</unstructured_citation>
          </citation>
          <citation key="ref27">
            <unstructured_citation>T. T. Ünlü, A. L. Mahamat, M. Turan, Leveraging deep learning for drowning and swimming prevention. WSEAS Transactions on Information Science and Applications, Vol. 22, 2025, pp 234-244. https://doi.org/10.37394/23209.2025.22.20.</unstructured_citation>
          </citation>
        </citation_list>
      </journal_article>
    </journal>
  </body>
</doi_batch>