lingtools.util module¶
-
lingtools.util.get_words_from_deepstruc(deepstruc_doc, punctuation=False)[source]¶ Receives a 3d list (a tagged doc) and returns a flat list of words.
Parameters: - deepstruc_doc – 3D list of (word, tag, lemma) (see Deep structure pos-tagged document)
- punctuation – boolean. When True, includes punctuation in the list. Default: False
-
lingtools.util.get_sentences_from_deepstruc(deepstruc_doc)[source]¶ Receives a 3d list (a tagged doc) and returns a flat list of sentences
Parameters: deepstruc_doc – 3D list of (word, tag, lemma) (see Deep structure pos-tagged document)
-
lingtools.util.get_wc_from_deepstruc(deepstruc_doc, punctuation=False)[source]¶ Receives a 3d list (a deep-structure doc) and returns the word count
Parameters: - deepstruc_doc – 3D list of (word, tag, lemma) (see Deep structure pos-tagged document)
- punctuation – boolean. When True, includes punctuation in the count. Default: False
-
lingtools.util.get_sentence_count_from_deepstruc(deepstruc_doc)[source]¶ Receives a 3d list (a tagged doc) and returns the word count
Parameters: deepstruc_doc – 3D list of (word, tag, lemma) (see Deep structure pos-tagged document)
-
lingtools.util.is_doc_deepstruc(input_document)[source]¶ Checks if an input document is a deep-structure document. See Deep structure pos-tagged document)
Parameters: input_document – the document to check
-
lingtools.util.is_punctuation(token)[source]¶ Checks if a token is punctuation. Returns a boolean.
Parameters: token – input token
-
lingtools.util.remove_punctuation_chars(word, replacement=None, regular_expression=None)[source]¶ Replaces punctuation characters with ‘replacement’
Parameters: - word – input word
- replacement – character that will substitute punctuation. Default = None
- regular_expression – The (regex module) regular expression to look for. Default = ur”[p{Punct}`~]+”
-
lingtools.util.remove_special_chars(word, replacement=None, regular_expression=None)[source]¶ Replaces special characters with ‘replacement’
Parameters: - word – input word
- replacement – character that will substitute special chars. Default = None
- regular_expression – The (regex module) regular expression to look for. Default = ur”p{S}+|[p{Punct}`]+|p{M}+|p{N}+”
-
lingtools.util.replace_digits_in_word(word, replacement=None)[source]¶ Replaces digits with another character.
Parameters: - word – the input word.
- replacement – a string. Default: #
-
lingtools.util.remove_punctuation(token)[source]¶ If a token consists of only punctuation character, returns an empty string.
Parameters: token – string
-
lingtools.util.transliterate_special_chars(document)[source]¶ Transliterates special characters to ASCII using unidecode (https://pypi.org/project/Unidecode/). Euro (€) and Pound (£) symbols are mainained!
Parameters: document – input document