lingtools.tokenizer module¶
-
class
lingtools.tokenizer.LocalTokenizer(language=None, encoding=None, tagger=None, stanford_url=None, simple_tokenizer=None, remove_stopwords=True, stopwords=None, replace_digits=True, remove_punctuation=True, lowercase=False, use_lemmas=False, replace_not=True, config_file=None)[source]¶ Bases:
objectUtility class to tokenize and tag documents. By default uses NLTK, but it is possible to use also a running Stanford POS Tagger server.
Parameters: - language – Choose between ‘english’, ‘dutch’, ‘russian’. Default: ‘english’
- encoding – Incoming text encoding. Default: utf-8
- tagger – nltk or stanford. Default: nltk
- stanford_url – If using stanford tagger, address of the server. Default: http://localhost:9000
- simple_tokenizer – Use a simple white space tokenizer
- remove_stopwords – Default True
- stopwords – The list of stopwords to ignore. Default: NLTK’s list
- replace_digits – Replaces digits with #. Default: True
- remove_punctuation – Removes punctuation. Default: True
- replace_not – Replaces “n’t” with “not”. Default: True
- lowercase – Converts text to lowercase before processing. Default: False
- use_lemmas – Uses lemmatized version of words. Default: False
- config_file (string) – Configuration file to use (optional. See Configuration).
-
get_deep_structure(input_document, transliterate=True, paragraph_sep=None, clean_extra_spaces=False)[source]¶ Takes a raw text and converts into a deep-structure document (paragraphs, sentences, tagged and lemmatized tokens).
Parameters: - input_document – Raw document (unicode)
- lowercase – boolean. If true, words will be converted to lowercase. Default: False
- remove_punctuation – boolean. If true, punctuation will be removed. Default: False
- transliterate – boolean. If true, special characters will be transliterated to ascii version. Default: True
- paragraph_sep – The paragraph boundary to split a document into paragraphs. Default: “nn”
-
simplify_deepstruc(deepstruc_doc, flatten=False)[source]¶ Convert the input deep-structure document into a 3D list of tokens
Parameters: deepstruc_doc – 3D list of (word, tag, lemma) (see Deep structure pos-tagged document)