lingtools.pybiber module¶
-
class
lingtools.pybiber.PyBiber(tokenizer_obj=None, config_file=None)[source]¶ Bases:
objectThese are the 67 features selected by Douglas Biber to reflect the linguistic structure of text.
Features are normalized by word count and multiplied by 1000 (range: 0-1000), with the exception of type_token_ratio and word_length.
- past_tense
- perfect_aspect_verbs
- present_tense
- place_adverbials
- time_adverbials
- first_person_pronouns
- second_person_pronouns
- third_person_pronouns
- pronoun_it
- demonstrative_pronouns
- indefinite_pronouns
- do_as_proverb
- wh_questions
- nominalizations
- gerunds
- nouns
- agentless_passives
- by_passives
- be_as_main_verb
- existential_there
- that_verb_complements
- that_adj_complements
- wh_clauses
- infinitives
- present_participial_clauses
- past_participial_clauses
- past_prt_whiz_deletions
- present_prt_whiz_deletions
- that_relatives_subj_position
- that_relatives_obj_position
- wh_relatives_subj_position
- wh_relatives_obj_position
- wh_relatives_pied_pipes
- sentence_relatives
- adv_subordinator_cause
- adv_sub_concesssion
- adv_sub_condition
- adv_sub_other
- prepositions
- attributive_adjectives
- predicative_adjectives
- adverbs
- type_token_ratio
- word_length
- conjuncts
- downtoners
- hedges
- amplifiers
- empathics
- discourse_particles
- demonstratives
- possibility_modals
- necessity_modals
- predictive_modals
- public_verbs
- private_verbs
- suasive_verbs
- seems_appear
- contractions
- that_deletion
- stranded_prepositions
- split_infinitives
- split_auxilaries
- phrasal_coordination
- non_phrasal_coordination
- synthetic_negation
- analytic_negation
Reference: Biber, D. (1988). Variation across speech and writing. Cambridge: Cambridge University Press. doi: 10.1017/CBO9780511621024
Parameters: - tokenizer_obj (LocalTokenizer) – initialized tokenizer object
- config_file (string) – Configuration file to use (optional. See Configuration).
-
static
get_feature_names()[source]¶ Returns a list of feature names in the same order as the features returned by
get_features().Returns: list of feature names Return type: list
-
get_features(deepstruc_doc)[source]¶ Returns features for one document, in the same order as returned by
get_feature_names().Parameters: deepstruc_doc (list) – The incoming document, pre-processed as a Deep structure pos-tagged document. Returns: list of features Return type: list
-
class
lingtools.pybiber.WordInfo(word, lemma, penn_tag, biber_tags=None)[source]¶ Bases:
objectWrapper class for words that include lemmas, POS-tags and Biber tags.
Any type of bracket tags – L/RRB for parenthesis (), L/RSB for square brackets [] and L/RCB for curly brackets {} – are converted to NLTK tags “(” and “)”, and the word and lemma are converted to round parenthesis too.
Parameters: - word – The original word
- lemma – The lemma
- penn_tag – the original penn tag
- biber_tags – the biber tags added to the word in the initial state, if any
-
activate(i)[source]¶ Activate Biber rule for this word.
Parameters: (int) (i) – the rule_id to activate
Display biber tags.
-
ends_with(ending)[source]¶ Checks if the word ends with ending
Parameters: (unicode) (ending) – the ending to check
Get all biber tags
-
has_biber_tag(biber_tag)[source]¶ Checks if this Wordobject has biber_tag
Parameters: biber_tag – tag to checl
-
has_lemma(lemma)[source]¶ Checks if the object has this lemma
Parameters: lemma – the lemma to verify
-
is_activated(i)[source]¶ Check if rule is activated for this word.
Parameters: (int) (i) – the rule_id to check