Introduction¶
This is a package that extracts linguistic features from texts.
Features extracted¶
These are the features extracted by the package:
Basic features¶
Some basic linguistic features, such as word count, sentence count, average sentence length, and incidence of certain grammatical classes.
Incidences are normalized by word count, so they range from 0 to 1.
- word_count
- sentence_count
- paragraph_count
- avg_word_length
- sd_word_length
- avg_sentence_length
- sd_sentence_length
- avg_paragraph_length
- sd_paragraph_length
- noun_incidence
- verb_incidence
- adjective_incidence
- adverb_incidence
- pronoun_incidence
- 1p_sing_pronoun_incidence
- 1p_pl_pronoun_incidence
- 2p_pronoun_incidence
- 3p_sing_pronoun_incidence
- 3p_pl_pronoun_incidence
Biber features¶
These are the 67 features selected by Douglas Biber [1] to reflect the linguistic structure of text.
Features are normalized by word count and multiplied by 1000 (range: 0-1000), with the exception of type_token_ratio and word_length.
- past_tense
- perfect_aspect_verbs
- present_tense
- place_adverbials
- time_adverbials
- first_person_pronouns
- second_person_pronouns
- third_person_pronouns
- pronoun_it
- demonstrative_pronouns
- indefinite_pronouns
- do_as_proverb
- wh_questions
- nominalizations
- gerunds
- nouns
- agentless_passives
- by_passives
- be_as_main_verb
- existential_there
- that_verb_complements
- that_adj_complements
- wh_clauses
- infinitives
- present_participial_clauses
- past_participial_clauses
- past_prt_whiz_deletions
- present_prt_whiz_deletions
- that_relatives_subj_position
- that_relatives_obj_position
- wh_relatives_subj_position
- wh_relatives_obj_position
- wh_relatives_pied_pipes
- sentence_relatives
- adv_subordinator_cause
- adv_sub_concesssion
- adv_sub_condition
- adv_sub_other
- prepositions
- attributive_adjectives
- predicative_adjectives
- adverbs
- type_token_ratio
- word_length
- conjuncts
- downtoners
- hedges
- amplifiers
- empathics
- discourse_particles
- demonstratives
- possibility_modals
- necessity_modals
- predictive_modals
- public_verbs
- private_verbs
- suasive_verbs
- seems_appear
- contractions
- that_deletion
- stranded_prepositions
- split_infinitives
- split_auxilaries
- phrasal_coordination
- non_phrasal_coordination
- synthetic_negation
- analytic_negation
MRC features¶
These features are the average and standard deviation scores of the words in the document according to the Medical Research Council (MRC) Psycholinguistic Database [2].
The ranges of the features are the same as informed in the MRC database. See http://websites.psychology.uwa.edu.au/school/MRCDatabase/mrc2.html for details.
- avg_Nlet (range: 1-23)
- sd_Nlet
- avg_Nphon (range: 0-19)
- sd_Nphon
- avg_Nsyl (range: 0-9)
- sd_Nsyl
- avg_K-F-freq (maximum frequency in file: 69971)
- sd_K-F-freq
- avg_K-F-ncats (maximum frequency in file: 69971)
- sd_K-F-ncats
- avg_K-F-nsamp (maximum frequency in file: 69971)
- sd_K-F-nsamp
- avg_T-L-freq
- sd_T-L-freq
- avg_Brown-freq (range of entries: 0 - 6833)
- sd_Brown-freq
- avg_Familiarity (range: 100 - 700)
- sd_Familiarity
- avg_Concreteness (range: 100 - 700)
- sd_Concreteness
- avg_Imageability (range: 100 - 700)
- sd_Imageability
- avg_Meaningfulness-Colorado (range: 100 - 700)
- sd_Meaningfulness-Colorado
- avg_Meaningfulness-Paivio (range: 100 - 700)
- sd_Meaningfulness-Paivio
- avg_Age-of-acquisition (range: 100 - 700)
- sd_Age-of-acquisition
ANEW features¶
These features are the average and standard deviation scores of the words in the document according to the Affective Norms for English Words (ANEW) [3].
- avg_valence (range: 1-9)
- sd_valence
- avg_arousal (range: 1-9)
- sd_arousal
- avg_dominance (range: 1-9)
- sd_dominance
SemanticVectors¶
These are measures of text coherence, using representations in the LSA vector space to calculate distances. The features are presented as average and standard deviations in the text.
Cosine distances can range from -1.0 to +1.0.
- avg_cosdis_adjacent_sentences
- sd_cosdis_adjacent_sentences
- avg_cosdis_all_sentences_in_paragraph
- sd_cosdis_all_sentences_in_paragraph
- avg_cosdis_adjacent_paragraphs
- sd_cosdis_adjacent_paragraphs
- avg_givenness_sentences
- sd_givenness_sentences
SentimentAnalysis¶
Sentiment Analysis scores using NLTK’s VADER sentiment analysis tool [4].
In the Vader lexicon, words are rated from -4 to +4 in the categories positive, negative and neutral. The compound value is a normalized combined score of the first three categories, ranging from -1.0 to +1.0.
- positive
- negative
- neutral
- compound
NER (Named Entity Recognition)¶
Recognized entities in the text, extracted using NLTK’s NER tool. The values are normalized by word count (range: 0-1).
- organization
- person
- gpe
- location
- facility
Extra features¶
Some extra basic features, all normalized by word count (range: 0-1).
To recognize English words, we use NLTK’s word and brown corpora.
- dollarsign
- eurosign
- poundsign
- numbers
- years
- englishwords
- nonenglishwords
References¶
- [1] Biber, D. (1988). Variation across speech and writing. Cambridge: Cambridge University Press. doi: 10.1017/CBO9780511621024
- [2] MRC Psycholinguistic Database. http://websites.psychology.uwa.edu.au/school/MRCDatabase/mrc2.html
- [3] Bradley, M. M., & Lang, P. J. (1999). Affective norms for English words (ANEW): Instruction manual and affective ratings (pp. 1-45). Technical report C-1, the center for research in psychophysiology, University of Florida.
- [4] Hutto, C.J. & Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.
Author¶
- Maira B. Carvalho (m.brandaocarvalho@uvt.nl), Tilburg University