lingtools.cosinedists module¶
-
class
lingtools.cosinedists.CosineDists(logging_interval=None, dictionary_file=None, vectorspace_file=None, vectorspace_name=None, tokenizer_obj=None, config_file=None)[source]¶ Bases:
objectCalculates matrices of cosine similarity values between documents.
Parameters: - logging_interval (int) – output logging info every logging_interval documents
- tokenizer_obj (LocalTokenizer) – initialized tokenizer object
- vectorspace_name (string) – Name of the vector space to be used (see Configuration).
- config_file (string) – Configuration file to use (optional. See Configuration).
-
get_matrix()[source]¶ Calculates a vector for each document (previously loaded with
load_list()orload_file()) and gets the cosine similarity between all documents.Returns: a NxN matrix (as a list of tuples) of cosine similarity values. Values range from -1.0 (most dissimilar) to 1.0 (most similar) Return type: list of tuples (docid1, docid2, cosine_similarity)
-
get_matrix_as_array()[source]¶ Calculates a vector for each document (previously loaded with
load_list()orload_file()) and gets the cosine similarity between all documents.Returns: a list with the field names and a NxN matrix (as a numpy array) of cosine similarity values. Values range from -1.0 (most dissimilar) to 1.0 (most similar) Return type: (field_names, array)
-
get_matrix_as_dict()[source]¶ Calculates a vector for each document (previously loaded with
load_list()orload_file()) and gets the cosine similarity between all documents.Returns: a NxN matrix (as a dictionary) of cosine similarity values. Values range from -1.0 (most dissimilar) to 1.0 (most similar) Return type: dict (matrix[docid1][docid2] = cosine_similarity)
-
get_nearest_neighbors(input_doc, n=5, algorithm=u'brute')[source]¶ Get the n nearest neighbor words for an input document.
Parameters: - input_doc (unicode) – the raw word or document
- n (int) – the number of nearest neighbors to retrieve. Default: 5
- algorithm (string) – which algorithm to use (see
sklearn.neighbors.NearestNeighborsfor possible values).
-
load_file(input_file, file_type=None, sep=None, encoding=None, skip=None, paragraph_sep=None, clean_extra_spaces=False)[source]¶ Loads a file for processing. See
lingtools.filereader.FileReader.load_file()for a description of the options.
-
load_list(doc_list, doc_index=None, paragraph_sep=None, clean_extra_spaces=False)[source]¶ Loads a list of (unicode) documents for processing. See
lingtools.filereader.FileReader.load_list()for a description of the options.Note that the list of documents loaded here is ignored if a file has been loaded with
load_file().
-
unload_file()[source]¶ Clear a previously loaded file. This function needs to be called if a file has been loaded before, but the user wants to process a list of documents (loaded with
load_list()) instead.