lingtools.semanticvectors module

class lingtools.semanticvectors.SemanticVectorsFeatures(dictionary_file=None, vectorspace_file=None, vectorspace_name=None, logging_interval=None, tokenizer_obj=None, config_file=None)[source]

Bases: object

These are measures of text coherence, using representations in a vector space to calculate cosine similarities. The features are presented as average and standard deviations in the text.

Cosine similarities can range from -1.0 to +1.0, where higher values indicate most similar documents.

  • avg_cosdis_adjacent_sentences
  • sd_cosdis_adjacent_sentences
  • avg_cosdis_all_sentences_in_paragraph
  • sd_cosdis_all_sentences_in_paragraph
  • avg_cosdis_adjacent_paragraphs
  • sd_cosdis_adjacent_paragraphs
  • avg_givenness_sentences
  • sd_givenness_sentences
Parameters:
  • logging_interval (int) – output logging info every logging_interval documents
  • tokenizer_obj (LocalTokenizer) – initialized tokenizer object
  • vectorspace_name (string) – Name of the vector space to be used (see Configuration).
  • config_file (string) – Configuration file to use (optional. See Configuration).
static get_feature_names()[source]

Returns a list of feature names in the same order as the features returned by get_features().

Returns:list of feature names
Return type:list
get_features(deepstruc_doc)[source]

Returns features for one document, in the same order as returned by get_feature_names().

Parameters:deepstruc_doc (list) – The incoming document, pre-processed as a Deep structure pos-tagged document.
Returns:list of features
Return type:list
static get_group_name()[source]

Get the human-readable name of the feature set.

Returns:unique identifier of the feature set
Return type:string