lingtools.dynamicthemes module¶
-
class
lingtools.dynamicthemes.DynamicThemes(logging_interval=None, vectorspace_name=None, tokenizer_obj=None, config_file=None)[source]¶ Bases:
objectCalculates cosine similarity values between a dynamically constructed theme and a list of documents.
Parameters: - logging_interval (int) – output logging info every logging_interval documents
- tokenizer_obj (LocalTokenizer) – initialized tokenizer object
- vectorspace_name (string) – Name of the vector space to be used (see Configuration).
- config_file (string) – Configuration file to use (optional. See Configuration).
-
get_theme_words()[source]¶ Get a clean list of the theme words. Only returns words that are recognized in the dictionary.
Returns: A list of words that will be considered in the theme. Return type: list
-
load_file(input_file, file_type=None, sep=None, encoding=None, skip=None, paragraph_sep=None, clean_extra_spaces=False)[source]¶ Loads a file for processing. See
lingtools.filereader.FileReader.load_file()for a description of the options.
-
load_list(doc_list, doc_index=None, paragraph_sep=None, clean_extra_spaces=False)[source]¶ Loads a list of (unicode) documents for processing. See
lingtools.filereader.FileReader.load_list()for a description of the options.Note that the list of documents loaded here is ignored if a file has been loaded with
load_file().
-
load_theme(theme_name, informed_theme_words, add_theme_name_to_list=False, remove_duplicates=True)[source]¶ Load a theme in the object so that the documents can be processed.
Note that words will NOT be preprocessed. So, for example, if the vector space was trained using uses lemmas, theme words need to be informed already as lemmas.
Parameters: - theme_name (unicode) – A name for the theme.
- informed_theme_words (list) – A list of theme words (as unicode). Compound words are allowed.
- add_theme_name_to_list (boolean) – Should the theme name be considered a theme word? Default: False
- remove_duplicates (boolean) – Should duplicate theme words be removed? Default: True
-
process()[source]¶ Processes previously loaded file or documents on the fly.
Calculates a vector for a user-given theme and gets the cosine similarity value between the document and the theme as a whole, and to each theme word separately.
The results are tuples in the following format: (theme_word, cosine_similarity, normalized_count). The first item refers to all theme words combined, while the following items list the similarity and count for each theme word separately.
Yields: tuple (docid,(results))
-
process_semantic_vectors()[source]¶ Process previously loaded documents on the fly.
Returns the representation of the document in the vector space.
Yields: a tuple (docid,(semantic_vector))
-
unload_file()[source]¶ Clear a previously loaded file. This function needs to be called if a file has been loaded before, but the user wants to process a list of documents (loaded with
load_list()) instead.