Deep Learning for Multi-Label Text Classification

This repository is my research project, and it is also a study of TensorFlow, Deep Learning (Fasttext, CNN, LSTM, etc.).

The main objective of the project is to solve the multi-label text classification problem based on Deep Neural Networks. Thus, the format of the data label is like [0, 1, 0, ..., 1, 1] according to the characteristics of such a problem.

Requirements

Python 3.6
Tensorflow 1.1 +
Numpy
Gensim

Innovation

Data part

Make the data support Chinese and English (Which use jieba seems easy).
Can use your own pre-trained word vectors (Which use gensim seems easy).
Add embedding visualization based on the tensorboard.

Model part

Add the correct L2 loss calculation operation.
Add gradients clip operation to prevent gradient explosion.
Add learning rate decay with exponential decay.
Add a new Highway Layer (Which is useful according to the model performance).
Add Batch Normalization Layer.

Code part

Can choose to train the model directly or restore the model from the checkpoint in train.py.
Can predict the labels via threshold and top-K in train.py and test.py.
Can calculate the evaluation metrics --- AUC & AUPRC.
Add test.py, the model test code, it can show the predicted values and predicted labels of the data in Testset when creating the final prediction file.
Add other useful data preprocess functions in data_helpers.py.
Use logging for helping to record the whole info (including parameters display, model training info, etc.).
Provide the ability to save the best n checkpoints in checkmate.py, whereas the tf.train.Saver can only save the last n checkpoints.

Data

See data format in data folder which including the data sample files.

Text Segment

You can use jieba package if you are going to deal with the Chinese text data.

Data Format

This repository can be used in other datasets (text classification) in two ways:

Modify your datasets into the same format of the sample.
Modify the data preprocess code in data_helpers.py.

Anyway, it should depend on what your data and task are.

🤔Before you open the new issue about the data format, please check the data_sample.json and read the other open issues first, because someone maybe ask me the same question already. For example: