lingtools.util module

lingtools.util.get_words_from_deepstruc(deepstruc_doc, punctuation=False)[source]

Receives a 3d list (a tagged doc) and returns a flat list of words.

Parameters:
  • deepstruc_doc – 3D list of (word, tag, lemma) (see Deep structure pos-tagged document)
  • punctuation – boolean. When True, includes punctuation in the list. Default: False
lingtools.util.get_sentences_from_deepstruc(deepstruc_doc)[source]

Receives a 3d list (a tagged doc) and returns a flat list of sentences

Parameters:deepstruc_doc – 3D list of (word, tag, lemma) (see Deep structure pos-tagged document)
lingtools.util.get_wc_from_deepstruc(deepstruc_doc, punctuation=False)[source]

Receives a 3d list (a deep-structure doc) and returns the word count

Parameters:
  • deepstruc_doc – 3D list of (word, tag, lemma) (see Deep structure pos-tagged document)
  • punctuation – boolean. When True, includes punctuation in the count. Default: False
lingtools.util.get_sentence_count_from_deepstruc(deepstruc_doc)[source]

Receives a 3d list (a tagged doc) and returns the word count

Parameters:deepstruc_doc – 3D list of (word, tag, lemma) (see Deep structure pos-tagged document)
lingtools.util.is_doc_deepstruc(input_document)[source]

Checks if an input document is a deep-structure document. See Deep structure pos-tagged document)

Parameters:input_document – the document to check
lingtools.util.is_punctuation(token)[source]

Checks if a token is punctuation. Returns a boolean.

Parameters:token – input token
lingtools.util.remove_punctuation_chars(word, replacement=None, regular_expression=None)[source]

Replaces punctuation characters with ‘replacement’

Parameters:
  • word – input word
  • replacement – character that will substitute punctuation. Default = None
  • regular_expression – The (regex module) regular expression to look for. Default = ur”[p{Punct}`~]+”
lingtools.util.remove_special_chars(word, replacement=None, regular_expression=None)[source]

Replaces special characters with ‘replacement’

Parameters:
  • word – input word
  • replacement – character that will substitute special chars. Default = None
  • regular_expression – The (regex module) regular expression to look for. Default = ur”p{S}+|[p{Punct}`]+|p{M}+|p{N}+”
lingtools.util.replace_digits_in_word(word, replacement=None)[source]

Replaces digits with another character.

Parameters:
  • word – the input word.
  • replacement – a string. Default: #
lingtools.util.remove_punctuation(token)[source]

If a token consists of only punctuation character, returns an empty string.

Parameters:token – string
lingtools.util.transliterate_special_chars(document)[source]

Transliterates special characters to ASCII using unidecode (https://pypi.org/project/Unidecode/). Euro (€) and Pound (£) symbols are mainained!

Parameters:document – input document
lingtools.util.memory_limit(config_file=None)[source]

Limit memory usage according to limit defined in the config file. See Memory.