AMLD Africa Workshop, September 04, 2021

Organizers

Khalil Mrini, Imane Khaouja, Ihsane gryech, Anass Sedrati and Abdelhak Mahmoudi

Moroccan Darija Wikipedia: Basics of Natural Language Processing for a Low-Resource Language

Description

NLP is a field that is in high demand, and where research progresses actively and quickly. Whereas language technology for languages like English and French is highly developed, low-resource languages (like most African indigenous languages) have been left behind and marginalized. There are many opportunities to create new tools for languages with few resources. In this tutorial, we take the example of Moroccan Darija, the national vernacular in Morocco. Our use case dataset will be the Moroccan Darija Wikipedia.

The participants will first learn statistical tools to analyze language in the tutorial. The tutorial will go over NLP notions including text pre-processing and tokenization, n-gram language modeling, n-gram frequency, topic modeling, and word embeddings. The tutorial consists of theoretical definitions and concrete examples in Python. The participants can then move to the practice part of the workshop, in teams of 1 to 5 people. Each team will be given the Moroccan Darija Wikipedia and will work on analyzing the dataset from an angle of their choice. At the end of the workshop, the teams will be invited to show their findings in a short presentation.

KhalilMrini/AMLD

AMLD Africa Workshop, September 04, 2021

Organizers

Moroccan Darija Wikipedia: Basics of Natural Language Processing for a Low-Resource Language

Description

Labs

Lab1: Wikipedia Darija Cleaning

Lab2: Wikipedia Darija Topic Detection

Lab3: NLP Tasks and Tools

Slides