Some corpus linguistic !
Some corpus linguistic ! Tonight! Lets make some corpus linguistic ! Corpus sources Once again, my years of geekiness and friendlyness helps! From Wiktionary contributions > Source > Authors blog > polite message > I got a link to a recent study and its gathered data. They already extracted the content from OpenSubtitles.org, watch yourself : http://opus.lingfil.uu.se/OpenSubtitles_v2.php [1] dig in it, its simply great. Counting and sorting From there on I needed some guiding for treatments. After Edouard Lopez gave me some directions and keywords : AWK , SHELL rather than pure REGEX, I went ahead to write down clear, concise, focused questions on StackOverflow[2][3] Given a multilingual .txt files such as: But where is Esope the holly Bastard! But where is ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? Console for Awk and Shell For speed reasons, we will work in console. The first thing to do is thus to tell the console in which folder are your files : cd /my/fo...