Arabic ACL corpus
- Title:
- Arabic ACL corpus
- Creator:
- Salah Elfahal Elebaed, Hoyam, Kasbi, Mohammed, Nasri, Mohammed, and Bouzoubaa, Karim
- Contributor:
- No@@No@@No@@ownFunds@@
- Publisher:
- International Journal of Computer Science Trends and Technology (IJCST)
- Identifier:
- http://hdl.handle.net/11372/LRT-4768
- Subject:
- Controlled Natural Language, Arabic CNL, ACL, Arabic Corpus, and and TEI.
- Type:
- text and corpus
- Description:
- This corpus constitutes all sentences representing the Arabic Controlled Language (ACL). It contains 551 sentences taken from four textbooks and websites dedicated to teach Arabic language to kids such as: a) First grade book, Republic of Sudan (كتاب الصف الاول جمهورية السودان), b) Al Jazeera Educational Site (موقع الجزيرة التعليمي), c) Bella Preparatory School Girls Forum (منتدى مدرسة بيلا الاعدادية بنات), and d) Albahr website (موقع انا البحر). These sentences are respecting 52 ACL rules. The average number of sentences for each rule is 10.6. All sentences in the corpus were analyzed by Farasa syntactic parser to confirm they are correctly analyzed. The validity of the parsing was done manually by linguist experts. The structure of this corpus is made of a header and a body. The header consists of a set of metadata that describe the corpus, such as the corpus name, the authors, the sources and further meta data. While the header is made of metadata, the body contains rules. Each rule has a code, a structure and all sentences respecting that rule. For each sentence, we store an id, the vowelledand unvowelled text as well as the result of parsing using Farasa.
- Language:
- Arabic
- Rights:
- Not specified
- Relation:
- http://arabic.emi.ac.ma/alelm/?q=Resources
http://www.ijcstjournal.org/volume-9/issue-6/IJCST-V9I6P8.pdf - Source:
- http://arabic.emi.ac.ma/alelm/?q=Resources
- Harvested from:
- LINDAT/CLARIAH-CZ repository
- Metadata only:
- true
- Date:
- 2021
The item or associated files might be "in copyright"; review the provided rights metadata:
- Not specified