|
Phrase extractor v.1.2
Find collocations in a text/corpus with MI | See also related N-Gram (lexical bundles) and Phrase Profiler |
Collocations are high mutual-information (high MI) phrases that are more frequent qua phrases than the average frequency of their component words separately. For ex, in some corpus Puerto Rico has appears 10 times, Puerto 11 times, and Rico 11 times. IE, these words appear mainly in this one phrase, not independently. The frequency of Puerto Rico is 10, averaged frequency of the component terms is (11+11)/2 = 10.5, so the ratio of phrase to individual word frequency is 10:10.5, or .95. MI becomes interesting at ratios of about .5 or .7 (see examples in the samples provided). Input Max 800k words.NOTE: Calculations are made entirely from within the entry text itself, not using a list from outside the text. For a list-based phrase/collocation profiler see PHRASE PROFILER