lingtools.mrc module¶
-
class
lingtools.mrc.MRCFeatures(tokenizer_obj=None, config_file=None)[source]¶ Bases:
objectExtract average and standard deviation scores of the words in the document according to the Medical Research Council (MRC) Psycholinguistic Database.
The ranges of the features are the same as informed in the MRC database.
- avg_Nlet (range: 1-23)
- sd_Nlet
- avg_Nphon (range: 0-19)
- sd_Nphon
- avg_Nsyl (range: 0-9)
- sd_Nsyl
- avg_K-F-freq (maximum frequency in file: 69971)
- sd_K-F-freq
- avg_K-F-ncats (maximum frequency in file: 69971)
- sd_K-F-ncats
- avg_K-F-nsamp (maximum frequency in file: 69971)
- sd_K-F-nsamp
- avg_T-L-freq
- sd_T-L-freq
- avg_Brown-freq (range of entries: 0 - 6833)
- sd_Brown-freq
- avg_Familiarity (range: 100 - 700)
- sd_Familiarity
- avg_Concreteness (range: 100 - 700)
- sd_Concreteness
- avg_Imageability (range: 100 - 700)
- sd_Imageability
- avg_Meaningfulness-Colorado (range: 100 - 700)
- sd_Meaningfulness-Colorado
- avg_Meaningfulness-Paivio (range: 100 - 700)
- sd_Meaningfulness-Paivio
- avg_Age-of-acquisition (range: 100 - 700)
- sd_Age-of-acquisition
Based on code from https://github.com/chbrown/lexicons
Reference: MRC Psycholinguistic Database. http://websites.psychology.uwa.edu.au/school/MRCDatabase/mrc2.html
Parameters: - tokenizer_obj (LocalTokenizer) – initialized tokenizer object
- config_file (string) – Configuration file to use (optional. See Configuration).
-
static
get_feature_names()[source]¶ Returns a list of feature names in the same order as the features returned by
get_features().Returns: list of feature names Return type: list
-
get_features(deepstruc_doc)[source]¶ Returns features for one document, in the same order as returned by
get_feature_names().Parameters: deepstruc_doc (list) – The incoming document, pre-processed as a Deep structure pos-tagged document. Returns: list of features Return type: list