lingtools.extra module¶
-
class
lingtools.extra.ExtraFeatures(tokenizer_obj=None, config_file=None)[source]¶ Bases:
objectSome extra basic features (numbers, special characters, English vs. non-English words), all normalized by word count (range: 0-1).
To recognize English words, we use NLTK’s word and brown corpora.
- dollarsign
- eurosign
- poundsign
- numbers
- years
- englishwords
- nonenglishwords
Parameters: - tokenizer_obj (LocalTokenizer) – initialized tokenizer object
- config_file (string) – Configuration file to use (optional. See Configuration).
-
static
get_feature_names()[source]¶ Returns a list of feature names in the same order as the features returned by
get_features().Returns: list of feature names Return type: list
-
get_features(deepstruc_doc)[source]¶ Returns features for one document, in the same order as returned by
get_feature_names().Parameters: deepstruc_doc (list) – The incoming document, pre-processed as a Deep structure pos-tagged document. Returns: list of features Return type: list