Charset encoding

Question

Charset encoding

GoogleCodeExporter opened this issue 8 years ago · 1 comments

GoogleCodeExporter commented 8 years ago

Ran and compiled on Ubuntu 14.04.

The input file is encoded in UTF-8 but output files (vocabulary and text 
vectors file) are encoded in ISO-8859-1.

All accents are wrong.

Original issue reported on code.google.com by pierpaol...@gmail.com on 10 Sep 2014 at 5:33

Answer 1 · 2016-04-06T17:17:55.000Z

confirmed on same os with python3

Original comment by jonatha...@gmail.com on 13 Mar 2015 at 1:13