Days ago, I wrote an article entitled creating Chinese/Japanese word clouds in Python. The article was written for a friend of mine who is learning the language for his research in mathematics and mathematical biology. By writing codes from scratch, one can learn data structures and time complexities. Hash tables and their worst case O(log n), binary search trees O(n) or whatever. Not that there is anything wrong with that. It just takes some time to think and write codes.
Instead of writing scripts, I usually get the same outputs from the GNU/Linux command line tools. As you may know, there are many other ways to do the same thing. And it is always a good idea to find easier, faster and/or more efficient ways. You do not always need to write your own code to get what you want.
Counting occurrences of Chinese nouns annotated by Stanford part-of-speech tagger, for instance, can be carried out by typing the following command.
is an option to remove localized settings that affect the sorting and comparison results
Despite the fact that Chinese text is written in non-alphanumeric, multi-byte characters, you can still take advantage of the major functions of UNIX and UNIX-like operating systems.
コメント
まだコメントはありません。最初のコメントをどうぞ。
コメントを残す