lingtools.wordvectorizer module¶
-
class
lingtools.wordvectorizer.WordVectorizer(dictionary_file=None, vectorspace_file=None, vectorspace_name=None, tokenizer_obj=None, config_file=None)[source]¶ Bases:
objectThe base class for the other feature extraction classes that rely on word embeddings.
Parameters: - dictionary_file (string) – Path for the dictionary for the vector space, which can be a text file with one word per line or a gensim dictionary
- vectorspace_file (string) – Path for a numpy matrix of N-dimensional vectors, in which each row represents a word. There must be a 1:1 match between the vectors in this file and the words in the dictionary file, in correct order.
- tokenizer_obj (LocalTokenizer) – initialized tokenizer object
- config_file (string) – Configuration file to use (optional. See Configuration).
-
calculate_cosine_similarity(vec1, vec2)[source]¶ Calculates a cosine similarity value between two vectors. Vectors must have the same dimensionality.
Parameters: - vec1 (numpy.array or list that can be converted to numpy.array) – first vector
- vec2 (numpy.array or list that can be converted to numpy.array) – second vector
-
combine_vectors(list_of_vectors, normalize=False)[source]¶ Combines vectors into one using vector addition, with optional normalizing to a unit vector.
Parameters: - list_of_vectors – the list of vectors to combine
- normalize (boolean) – Should the vectors be normalized to a unit vector?
-
get_indices_from_tokens(list_of_tokens)[source]¶ Receives a flat list of tokens and returns their indices in the dictionary
Parameters: list_of_tokens (flat list) – list of tokens to process
-
get_structured_vectors(structure_of_tokens)[source]¶ Converts a 3-dimensional list of tokens (document > paragraph > sentences) into a 3-dimensional list of vectors
Parameters: structure_of_tokens (list) – 3-dimensional list of tokens
-
get_vector_from_raw_doc(raw_doc, normalize=False)[source]¶ Receives a raw document and returns the vectorized version of the document.
Parameters: - raw_doc (unicode) – raw document
- normalize (boolean) – Should the vectors be normalized to a unit vector?