A sentence tokenizer NLP tool for the Tamil language
Command-line utility to perform sentence tokenization on a given Tamil corpus text file.
python sentence_tokenizer.py <input_file>
python sentence_tokenizer.py -h
- No preprocessing needed
- Works on any OS which supports Python 3
- Handles input file of any size