lingtools.cosinedists module

class lingtools.cosinedists.CosineDists(logging_interval=None, dictionary_file=None, vectorspace_file=None, vectorspace_name=None, tokenizer_obj=None, config_file=None)[source]

Bases: object

Calculates matrices of cosine similarity values between documents.

Parameters:
  • logging_interval (int) – output logging info every logging_interval documents
  • tokenizer_obj (LocalTokenizer) – initialized tokenizer object
  • vectorspace_name (string) – Name of the vector space to be used (see Configuration).
  • config_file (string) – Configuration file to use (optional. See Configuration).
get_matrix()[source]

Calculates a vector for each document (previously loaded with load_list() or load_file()) and gets the cosine similarity between all documents.

Returns:a NxN matrix (as a list of tuples) of cosine similarity values. Values range from -1.0 (most dissimilar) to 1.0 (most similar)
Return type:list of tuples (docid1, docid2, cosine_similarity)
get_matrix_as_array()[source]

Calculates a vector for each document (previously loaded with load_list() or load_file()) and gets the cosine similarity between all documents.

Returns:a list with the field names and a NxN matrix (as a numpy array) of cosine similarity values. Values range from -1.0 (most dissimilar) to 1.0 (most similar)
Return type:(field_names, array)
get_matrix_as_dict()[source]

Calculates a vector for each document (previously loaded with load_list() or load_file()) and gets the cosine similarity between all documents.

Returns:a NxN matrix (as a dictionary) of cosine similarity values. Values range from -1.0 (most dissimilar) to 1.0 (most similar)
Return type:dict (matrix[docid1][docid2] = cosine_similarity)
get_nearest_neighbors(input_doc, n=5, algorithm=u'brute')[source]

Get the n nearest neighbor words for an input document.

Parameters:
  • input_doc (unicode) – the raw word or document
  • n (int) – the number of nearest neighbors to retrieve. Default: 5
  • algorithm (string) – which algorithm to use (see sklearn.neighbors.NearestNeighbors for possible values).
load_file(input_file, file_type=None, sep=None, encoding=None, skip=None, paragraph_sep=None, clean_extra_spaces=False)[source]

Loads a file for processing. See lingtools.filereader.FileReader.load_file() for a description of the options.

load_list(doc_list, doc_index=None, paragraph_sep=None, clean_extra_spaces=False)[source]

Loads a list of (unicode) documents for processing. See lingtools.filereader.FileReader.load_list() for a description of the options.

Note that the list of documents loaded here is ignored if a file has been loaded with load_file().

unload_file()[source]

Clear a previously loaded file. This function needs to be called if a file has been loaded before, but the user wants to process a list of documents (loaded with load_list()) instead.