lingtools.freqanalyzer module¶
-
class
lingtools.freqanalyzer.FreqAnalyzer(logging_interval=None, tokenizer_obj=None, config_file=None)[source]¶ Bases:
objectAnalyzes word frequencies to compare two groups of documents.
Parameters: - logging_interval (int) – output logging info every logging_interval documents
- tokenizer_obj (LocalTokenizer) – initialized tokenizer object
- config_file (string) – Configuration file to use (optional. See Configuration).
-
create_termdocmatrix()[source]¶ Creates a term-document matrix using the documents in all the files loaded with
load_file().This function must be called before
get_distinctive_words().
-
get_distinctive_words(group1, group2, alpha=0.05)[source]¶ Compares two groups of documents and extracts the most distinctive words between them using Chi-square.
Before calling this function, you need to load at least two files with
load_file()and callcreate_termdocmatrix().Parameters: - group1 (string) – the name of the first group
- group2 (string) – the name of the second group
Returns: a tuple (word, chi-square, pval, word rate (per 1000 words) in group 1, word rate (per 1000 words) in group 2, group where word appears most)
-
load_file(input_file, group_name=None, file_type=None, paragraph_sep=None, sep=None, encoding=None, skip=None, clean_extra_spaces=False)[source]¶ Loads a file for processing.
Parameters: - input_file – path to the file
- group_name (string) – a group name (as an identifier) for the documents in the file
- type – txt_single|txt_multiple|csv. Default: txt_single
- skip – How many lines to skip (in txt_multiple and csv files). Default: 0
- paragraph_sep – New paragraph separator (in txt_single and zip files). Default:
- sep – Separator in the CSV file. Default: ;
- clean_extra_spaces – boolean. If true, double spaces and extra spaces between new lines will be removed. Default: False
- encoding – Encoding of input file. Default: utf-8