lingtools.tokenizer module

class lingtools.tokenizer.LocalTokenizer(language=None, encoding=None, tagger=None, stanford_url=None, simple_tokenizer=None, remove_stopwords=True, stopwords=None, replace_digits=True, remove_punctuation=True, lowercase=False, use_lemmas=False, replace_not=True, config_file=None)[source]

Bases: object

Utility class to tokenize and tag documents. By default uses NLTK, but it is possible to use also a running Stanford POS Tagger server.

Parameters:
  • language – Choose between ‘english’, ‘dutch’, ‘russian’. Default: ‘english’
  • encoding – Incoming text encoding. Default: utf-8
  • tagger – nltk or stanford. Default: nltk
  • stanford_url – If using stanford tagger, address of the server. Default: http://localhost:9000
  • simple_tokenizer – Use a simple white space tokenizer
  • remove_stopwords – Default True
  • stopwords – The list of stopwords to ignore. Default: NLTK’s list
  • replace_digits – Replaces digits with #. Default: True
  • remove_punctuation – Removes punctuation. Default: True
  • replace_not – Replaces “n’t” with “not”. Default: True
  • lowercase – Converts text to lowercase before processing. Default: False
  • use_lemmas – Uses lemmatized version of words. Default: False
  • config_file (string) – Configuration file to use (optional. See Configuration).
get_deep_structure(input_document, transliterate=True, paragraph_sep=None, clean_extra_spaces=False)[source]

Takes a raw text and converts into a deep-structure document (paragraphs, sentences, tagged and lemmatized tokens).

Parameters:
  • input_document – Raw document (unicode)
  • lowercase – boolean. If true, words will be converted to lowercase. Default: False
  • remove_punctuation – boolean. If true, punctuation will be removed. Default: False
  • transliterate – boolean. If true, special characters will be transliterated to ascii version. Default: True
  • paragraph_sep – The paragraph boundary to split a document into paragraphs. Default: “nn”
simplify_deepstruc(deepstruc_doc, flatten=False)[source]

Convert the input deep-structure document into a 3D list of tokens

Parameters:deepstruc_doc – 3D list of (word, tag, lemma) (see Deep structure pos-tagged document)