tokeniser

There are 12 repositories under tokeniser topic.

dragonofmercy/Tokenize2
Tokenize2 is a plugin which allows your users to select multiple items from a predefined list or ajax, using autocompletion as they type to find each item. You may have seen a similar type of text entry when filling in the recipients field sending messages on facebook or tags on tumblr.
Language:JavaScript83 6 7725
LanguageMachines/ucto
Unicode tokeniser. Ucto tokenizes text files: it separates words from punctuation, and splits sentences. It offers several other basic preprocessing steps such as changing case that you can all use to make your text suited for further processing such as indexing, part-of-speech tagging, or machine translation. Ucto comes with tokenisation rules for several languages and can be easily extended to suit other languages. It has been incorporated for tokenizing Dutch text in Frog, our Dutch morpho-syntactic processor. http://ilk.uvt.nl/ucto --
Language:C++69 12 9314
andreihar/taibun
Taiwanese Hokkien Transliterator and Tokeniser
Language:Python33 1 101
jonsafari/tok-tok
A fast, simple, multilingual tokenizer
Language:Python29 5 13
ben-sb/jisu
JavaScript Parser
Language:TypeScript10 1 00
kuhumcst/rtfreader
Text segmenter and tokeniser for Danish, English and other languages. Reads an RTF or flat text file and outputs the text, one line per sentence & optionally tokenized.
Language:C++8 3 24
ztjhz/word-piece-tokenizer
A Lightweight Word Piece Tokenizer
Language:Python6 3 10
phughesmcr/happynodetokenizer
Javascript port of HappyFunTokenizer.py by Christopher Potts and HappierFunTokenizing.py by H. Andrew Schwartz
Language:TypeScript5 1 00
andreihar/taibun.js
Taiwanese Hokkien Transliterator and Tokeniser
Language:JavaScript2 1 00
TomHarte/bas2uef
Converts BBC BASIC 2 source code into UEF files.
Language:C++2 1 0
P0u4a/casio-parser
Find out useful information about your Casio watch
Language:Rust00
Abhigyan126/SentencePiece-Tokenisation
A python and rust implementation of SentencePiece (A language-independent subword tokeniser and de-tokeniser developed by Google)
Language:Rust1 0

tokeniser

dragonofmercy/Tokenize2

LanguageMachines/ucto

andreihar/taibun

jonsafari/tok-tok

ben-sb/jisu

kuhumcst/rtfreader

ztjhz/word-piece-tokenizer

phughesmcr/happynodetokenizer

andreihar/taibun.js

TomHarte/bas2uef

P0u4a/casio-parser

Abhigyan126/SentencePiece-Tokenisation