lingtools.freqanalyzer module

class lingtools.freqanalyzer.FreqAnalyzer(logging_interval=None, tokenizer_obj=None, config_file=None)[source]

Bases: object

Analyzes word frequencies to compare two groups of documents.

Parameters:
  • logging_interval (int) – output logging info every logging_interval documents
  • tokenizer_obj (LocalTokenizer) – initialized tokenizer object
  • config_file (string) – Configuration file to use (optional. See Configuration).
create_termdocmatrix()[source]

Creates a term-document matrix using the documents in all the files loaded with load_file().

This function must be called before get_distinctive_words().

get_distinctive_words(group1, group2, alpha=0.05)[source]

Compares two groups of documents and extracts the most distinctive words between them using Chi-square.

Before calling this function, you need to load at least two files with load_file() and call create_termdocmatrix().

Parameters:
  • group1 (string) – the name of the first group
  • group2 (string) – the name of the second group
Returns:

a tuple (word, chi-square, pval, word rate (per 1000 words) in group 1, word rate (per 1000 words) in group 2, group where word appears most)

load_file(input_file, group_name=None, file_type=None, paragraph_sep=None, sep=None, encoding=None, skip=None, clean_extra_spaces=False)[source]

Loads a file for processing.

Parameters:
  • input_file – path to the file
  • group_name (string) – a group name (as an identifier) for the documents in the file
  • type – txt_single|txt_multiple|csv. Default: txt_single
  • skip – How many lines to skip (in txt_multiple and csv files). Default: 0
  • paragraph_sep – New paragraph separator (in txt_single and zip files). Default:
  • sep – Separator in the CSV file. Default: ;
  • clean_extra_spaces – boolean. If true, double spaces and extra spaces between new lines will be removed. Default: False
  • encoding – Encoding of input file. Default: utf-8