Classifying Suspicious Content Using Frequency Analysis


This paper details an experiment to explore the use of chi by degrees of freedom (CBDF) and Log-Likelihood statistical similarity measures with single word and bigram frequencies as a means of discriminating subject content in order to classify samples of chat texts as dangerous, suspicious or innocent. The control for these comparisons was a set of… (More)

9 Figures and Tables


