lingtools.sentimentanalysis module

class lingtools.sentimentanalysis.SentimentAnalysisFeatures(tokenizer_obj=None, config_file=None)[source]

Bases: object

Sentiment Analysis scores using NLTK’s VADER sentiment analysis tool.

In the Vader lexicon, words are rated from -4 to +4 in the categories positive, negative and neutral. The compound value is a normalized combined score of the first three categories, ranging from -1.0 to +1.0.

  • positive
  • negative
  • neutral
  • compound

Hutto, C.J. & Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.

Parameters:
  • tokenizer_obj (LocalTokenizer) – initialized tokenizer object
  • config_file (string) – Configuration file to use (optional. See Configuration).
static get_feature_names()[source]

Returns a list of feature names in the same order as the features returned by get_features().

Returns:list of feature names
Return type:list
get_features(deepstruc_doc)[source]

Returns features for one document, in the same order as returned by get_feature_names().

Parameters:deepstruc_doc (list) – The incoming document, pre-processed as a Deep structure pos-tagged document.
Returns:list of features
Return type:list
static get_group_name()[source]

Get the human-readable name of the feature set.

Returns:unique identifier of the feature set
Return type:string