lingtools.pybiber module

class lingtools.pybiber.PyBiber(tokenizer_obj=None, config_file=None)[source]

Bases: object

These are the 67 features selected by Douglas Biber to reflect the linguistic structure of text.

Features are normalized by word count and multiplied by 1000 (range: 0-1000), with the exception of type_token_ratio and word_length.

  • past_tense
  • perfect_aspect_verbs
  • present_tense
  • place_adverbials
  • time_adverbials
  • first_person_pronouns
  • second_person_pronouns
  • third_person_pronouns
  • pronoun_it
  • demonstrative_pronouns
  • indefinite_pronouns
  • do_as_proverb
  • wh_questions
  • nominalizations
  • gerunds
  • nouns
  • agentless_passives
  • by_passives
  • be_as_main_verb
  • existential_there
  • that_verb_complements
  • that_adj_complements
  • wh_clauses
  • infinitives
  • present_participial_clauses
  • past_participial_clauses
  • past_prt_whiz_deletions
  • present_prt_whiz_deletions
  • that_relatives_subj_position
  • that_relatives_obj_position
  • wh_relatives_subj_position
  • wh_relatives_obj_position
  • wh_relatives_pied_pipes
  • sentence_relatives
  • adv_subordinator_cause
  • adv_sub_concesssion
  • adv_sub_condition
  • adv_sub_other
  • prepositions
  • attributive_adjectives
  • predicative_adjectives
  • adverbs
  • type_token_ratio
  • word_length
  • conjuncts
  • downtoners
  • hedges
  • amplifiers
  • empathics
  • discourse_particles
  • demonstratives
  • possibility_modals
  • necessity_modals
  • predictive_modals
  • public_verbs
  • private_verbs
  • suasive_verbs
  • seems_appear
  • contractions
  • that_deletion
  • stranded_prepositions
  • split_infinitives
  • split_auxilaries
  • phrasal_coordination
  • non_phrasal_coordination
  • synthetic_negation
  • analytic_negation

Reference: Biber, D. (1988). Variation across speech and writing. Cambridge: Cambridge University Press. doi: 10.1017/CBO9780511621024

Parameters:
  • tokenizer_obj (LocalTokenizer) – initialized tokenizer object
  • config_file (string) – Configuration file to use (optional. See Configuration).
static get_feature_names()[source]

Returns a list of feature names in the same order as the features returned by get_features().

Returns:list of feature names
Return type:list
get_features(deepstruc_doc)[source]

Returns features for one document, in the same order as returned by get_feature_names().

Parameters:deepstruc_doc (list) – The incoming document, pre-processed as a Deep structure pos-tagged document.
Returns:list of features
Return type:list
static get_group_name()[source]

Get the human-readable name of the feature set.

Returns:unique identifier of the feature set
Return type:string
get_named_features(input_document)[source]

Deprecated.

update_rule_objects(rules_list, processed_doc)[source]

Iterates over the tokens in the processed document and applies all Biber rules, in the correct processing order.

Parameters:
  • rules_list –
  • processed_doc –
class lingtools.pybiber.WordInfo(word, lemma, penn_tag, biber_tags=None)[source]

Bases: object

Wrapper class for words that include lemmas, POS-tags and Biber tags.

Any type of bracket tags – L/RRB for parenthesis (), L/RSB for square brackets [] and L/RCB for curly brackets {} – are converted to NLTK tags “(” and “)”, and the word and lemma are converted to round parenthesis too.

Parameters:
  • word – The original word
  • lemma – The lemma
  • penn_tag – the original penn tag
  • biber_tags – the biber tags added to the word in the initial state, if any
activate(i)[source]

Activate Biber rule for this word.

Parameters:(int) (i) – the rule_id to activate
display_tags()[source]

Display biber tags.

ends_with(ending)[source]

Checks if the word ends with ending

Parameters:(unicode) (ending) – the ending to check
get_biber_tags()[source]

Get all biber tags

get_lemma()[source]

Get the word lemma

get_penn_tag()[source]

Get the word Penn tag

get_word()[source]

Get the word in the original form.

has_biber_tag(biber_tag)[source]

Checks if this Wordobject has biber_tag

Parameters:biber_tag – tag to checl
has_lemma(lemma)[source]

Checks if the object has this lemma

Parameters:lemma – the lemma to verify
is_activated(i)[source]

Check if rule is activated for this word.

Parameters:(int) (i) – the rule_id to check
set_biber_tag(biber_tag)[source]

Append biber tag to list of tags.

Parameters:biber_tag – the biber tag to append.
set_lemma(lemma)[source]

Set the lemma

Parameters:lemma – the word’s lemma
set_word(word)[source]

Set the word.

Parameters:word – original format of the word