lingtools.wordvectorizer module

class lingtools.wordvectorizer.WordVectorizer(dictionary_file=None, vectorspace_file=None, vectorspace_name=None, tokenizer_obj=None, config_file=None)[source]

Bases: object

The base class for the other feature extraction classes that rely on word embeddings.

Parameters:
  • dictionary_file (string) – Path for the dictionary for the vector space, which can be a text file with one word per line or a gensim dictionary
  • vectorspace_file (string) – Path for a numpy matrix of N-dimensional vectors, in which each row represents a word. There must be a 1:1 match between the vectors in this file and the words in the dictionary file, in correct order.
  • tokenizer_obj (LocalTokenizer) – initialized tokenizer object
  • config_file (string) – Configuration file to use (optional. See Configuration).
calculate_cosine_similarity(vec1, vec2)[source]

Calculates a cosine similarity value between two vectors. Vectors must have the same dimensionality.

Parameters:
  • vec1 (numpy.array or list that can be converted to numpy.array) – first vector
  • vec2 (numpy.array or list that can be converted to numpy.array) – second vector
combine_vectors(list_of_vectors, normalize=False)[source]

Combines vectors into one using vector addition, with optional normalizing to a unit vector.

Parameters:
  • list_of_vectors – the list of vectors to combine
  • normalize (boolean) – Should the vectors be normalized to a unit vector?
get_indices_from_tokens(list_of_tokens)[source]

Receives a flat list of tokens and returns their indices in the dictionary

Parameters:list_of_tokens (flat list) – list of tokens to process
get_structured_vectors(structure_of_tokens)[source]

Converts a 3-dimensional list of tokens (document > paragraph > sentences) into a 3-dimensional list of vectors

Parameters:structure_of_tokens (list) – 3-dimensional list of tokens
get_vector_from_raw_doc(raw_doc, normalize=False)[source]

Receives a raw document and returns the vectorized version of the document.

Parameters:
  • raw_doc (unicode) – raw document
  • normalize (boolean) – Should the vectors be normalized to a unit vector?
get_vector_from_tokens(list_of_tokens, normalize=False)[source]

Receives a list of tokens and returns the vectorized version of the document.

Parameters:
  • list_of_tokens (list) – flat list of tokens
  • normalize (boolean) – Should the vectors be normalized to a unit vector?
get_vectors_from_indices(list_of_indices)[source]

Gets the vectors from the indices.

Parameters:list_of_indices (list) – list of indices of words in the dictionary