POS-Taggging Malay Corpus: A Novel Approach Based on Maximum Entropy


  • Juhaida Abu Bakar
  • Khairuddin Khairuddin
  • Mohammad Faidzul Nasrudin
  • Mohd Zamri Murah






NLP pipeline task, POS-tags, tagging approach, Malay language, Jawi.


Jawi and Roman scripts are represented Malay language. In the past, Jawi writings are widely used by the Malay community and foreigners; and it can be seen in the old documents. Old documents face the risk of background damage. In order to preserve this valuable information, there are significant needs to automated Jawi materials. Based on previous literature, POS-tags are known as the first phase in the automated text analysis; and the development of language technologies can barely initiate without this phase. We highlight the existing POS-tags approaches; and suggest the development of Malay Jawi POS-tags using extended ME-based approach on NUWT Corpus. Results have shown that the proposed model yielded a higher accuracy in comparison to the state-of-the-art model.




