lingtools.basicfeatures module

class lingtools.basicfeatures.BasicFeatures(tokenizer_obj=None, config_file=None)[source]

Bases: object

Extracts some basic linguistic features, such as word count, sentence count, average sentence length, and incidence of certain grammatical classes.

Incidences are normalized by word count, so they range from 0 to 1.

  • word_count
  • sentence_count
  • paragraph_count
  • avg_word_length
  • sd_word_length
  • avg_sentence_length
  • sd_sentence_length
  • avg_paragraph_length
  • sd_paragraph_length
  • noun_incidence
  • verb_incidence
  • adjective_incidence
  • adverb_incidence
  • pronoun_incidence
  • 1p_sing_pronoun_incidence
  • 1p_pl_pronoun_incidence
  • 2p_pronoun_incidence
  • 3p_sing_pronoun_incidence
  • 3p_pl_pronoun_incidence
Parameters:
  • tokenizer_obj (LocalTokenizer) – initialized tokenizer object
  • config_file (string) – Configuration file to use (optional. See Configuration).
static get_feature_names()[source]

Returns a list of feature names in the same order as the features returned by get_features().

Returns:list of feature names
Return type:list
get_features(deepstruc_doc)[source]

Returns features for one document, in the same order as returned by get_feature_names().

Parameters:deepstruc_doc (list) – The incoming document, pre-processed as a Deep structure pos-tagged document.
Returns:list of features
Return type:list
static get_group_name()[source]

Get the human-readable name of the feature set.

Returns:unique identifier of the feature set
Return type:string