How to code with lingtools¶
Table of contents
Coding with lingtools¶
Lingtools is a package that provides the following functions:
- A tokenizer wrapper class, which is used consistently by the other classes (class
lingtools.tokenizer.LocalTokenizer). - A feature extractor class, which processes text into numerical features (class
lingtools.featureextractor.FeatureExtractor). This class can be extended with add-ons. - A class to calculate cosine similarity values and nearest neighbors (class
lingtools.cosinedists.CosineDists). - A class to evaluate documents using dynamically defined themes (class
lingtools.dynamicthemes.DynamicThemes). - A frequency analyzer class, which compares texts or groups of texts to extract most distinctive words (class
lingtools.freqanalyzer.FreqAnalyzer).
The lingtools package has other auxiliary classes, but the user does not need to know them to use the module. If you want to know more about the remaining classes, you can check the lingtools API.
Roughly speaking, to use lingtools, you need to perform the following steps:
- Instantiate a tokenizer object with the desired settings;
- Instantiate an object of the class providing the functionality you want (feature extractor, dynamic themes, etc) and loading the desired vector space;
- Load a file or a list of documents in the class;
- Call the appropriate processing function.
These steps will be described in detail below.
Local tokenizer¶
The class lingtools.tokenizer.LocalTokenizer ensures that the user has control over how the document is pre-processed and that the pre-processing steps are the same across different functionalities.
This class can use either NLTK’s (https://www.nltk.org/) pos-tagging and lemmatizing functions, or a server running Stanford CoreNLP (https://stanfordnlp.github.io/CoreNLP/). NLTK is the default option because it runs locally without any further configuration.
This is how you instantiate a tokenizer object using the default NLTK tagger:
from lingtools.tokenizer import LocalTokenizer
tokenizer_settings = {"language":"english",
"encoding": "utf-8",
"tagger": "nltk",
"stanford_url": None,
"simple_tokenizer": False,
"remove_stopwords": True,
"stopwords": None, # Use NLTK default stopwords
"replace_digits": True,
"remove_punctuation": True,
"replace_not": True,
"lowercase": False,
"use_lemmas": True}
tokenizer_obj = LocalTokenizer(**tokenizer_settings)
To use a Stanford CoreNLP server, it is necessary to inform the URL in the format <protocol>://<server>:<port> .
tokenizer_settings_stanford = {"language":"english",
"encoding": "utf-8",
"tagger": "stanford",
"stanford_url": "http://localhost:9000",
"simple_tokenizer": False,
"remove_stopwords": True,
"stopwords": None, # Use NLTK default stopwords
"replace_digits": True,
"remove_punctuation": True,
"replace_not": True,
"lowercase": False,
"use_lemmas": True}
tokenizer_stanford_obj = LocalTokenizer(**tokenizer_settings_stanford)
Make sure the CoreNLP server is configured to output documents in the ‘json’ format and that it uses the following annotators, in this order: ‘tokenize,pos,lemma,ssplit’.
The stanford-corenlp-full-2018-02-27 server can be started (locally, in port 9002, with 4Gb memory) in the following manner:
# Run the server using all jars in the current directory (e.g., the CoreNLP home directory)
java -mx4g -cp "*" edu.stanford.nlp.pipeline.StanfordCoreNLPServer -port 9002 -timeout 15000
For more information on how to run the CoreNLP server, see https://stanfordnlp.github.io/CoreNLP/corenlp-server.html.
Deep structure pos-tagged document¶
The LocalTokenizer class generates a specially formatted tagged document, the “deep structure pos-tagged document”. Rarely, if ever, this format will be used directly by the user when utilizing the lingtools package.
The “deep structure pos-tagged document” is a 3-dimensional list of PennBank POS-tagged and lemmatized tokens (list of paragraphs, which contains a list of sentences, which contains a list of tagged tokens). The tagged tokens have three components: the original token, its POS-tag and its lemma.
For example, take the following raw text (double line break indicates a new paragraph):
The sky is blue.
My name is John.
The LocalTokenizer processes this document into the following deep structure pos-tagged document:
[ # document level
[ # paragraph level
[ # sentence level
(u'The', 'DT', 'the'),
(u'sky', 'NN', u'sky'),
(u'is', 'VBZ', u'be'),
(u'blue', 'JJ', u'blue'),
(u'.', '.', u'.')]
],
[
[
(u'My', 'PRP$', u'my'),
(u'name', 'NN', u'name'),
(u'is', 'VBZ', u'be'),
(u'John', 'NNP', u'john'),
(u'.', '.', u'.')
]
]
]
Vector spaces¶
Lingtools requires a vector space to perform most of its functions. Lingtools comes with a Default vector space, which will be used if no other option is selected. Other vector spaces can be configured and called by name when instantiating a class object. See Vector spaces configuration for details.
Loading files or lists¶
For most lingtools classes, you have to first load the instantiated object with a file or word/document list that will be processed on the fly. The appropriate processing function can only be called once the file or the list of documents is loaded in the object.
Files¶
The load_file() function allows lingtools to process large files without having to load everything in memory first.
Lingtools accepts the following file formats:
- Multiple documents as comma separated values (.csv)
- Multiple documents as plain text (.txt)
- Single document as plain text (.txt)
For CSV files, it is possible to choose which separator character to use.
For multiple document files, it is possible to indicate how many lines to skip (to account for a header, for example).
For single document text files, it is possible to indicate which character should be consider a paragraph splitter.
For all options, it is possible to inform the encoding (default: utf-8).
Important
In CSV files, it is expected that the first column contains the ID of the document, and the second column contains the text. All remaining columns are ignored.
Lists¶
It is also possible to load simple lists of documents instead of files, using the function load_list().
Important
Always use unicode strings! Using byte strings will cause lingtools to throw an Exception. For more information, see https://docs.python.org/2/howto/unicode.html.
Extracting features¶
To extract features from text, you need to instantiate a feature extractor object from class FeatureExtractor, using a previously instantiated LocalTokenizer object (see Local tokenizer):
from lingtools.featureextractor import FeatureExtractor
# Create a dictionary with settings, including the tokenizer object created previously
feature_extractor_settings = {"feature_sets": ["anew", "biber", "mrc"], # Choose which modules to extract
"vectorspace_name": 'glove', # Non-default vector space (optional)
"tokenizer_obj": tokenizer_obj
}
# Instantiate feature extractor object
feature_extractor = FeatureExtractor(**feature_extractor_settings)
Then, you are ready to load a file for extracting the features using the function load_file():
feature_extractor.load_file('test.csv',
sep=';',
encoding='utf-8',
skip=1, # Skip the header
file_type='csv')
Alternatively, a list of strings can be loaded using the function load_list():
input_docs = [ u"This is the first document",
u"This is the second document" ]
# It is possible to load named indices for the documents
input_docs_indices = [u"doc1", u"doc2"]
feature_extractor.load_list(input_docs,
doc_index=input_docs_indices)
The loaded file is processed document by document by the function process():
# Process results
for idx, result in feature_extractor.process():
# Get a result tuple from the object:
# (group_id, group_code, feature_id, feature_code, feature_value)
print (idx, result.get_results_as_tuples())
An output file can be created with the help of the function get_feature_names(), which can be used to generate a header, and function process_simple(), which outputs a simple list of values (instead of a whole feature container):
# Create a header first
with open('result.csv', 'w') as f:
header_items = ['docid'] + feature_extractor.get_feature_names()
header_line = ";".join(header_items)
f.write(header_line+"\n")
# Iterate over the documents and append values to the output file
for idx, result_list in feature_extractor.process_simple():
# Format the line using the feature value
result_items = [idx] + [str(x) for x in result_list]
result_line = ";".join(result_items)
# Write the result
with open('result.csv', 'a') as f:
f.write(result_line + "\n")
Add-ons¶
The FeatureExtractor object accepts add-on modules. Add-ons are classes that implement at least the following methods:
- get_group_name(): returns a name for the feature set, so that it can be identified among the other groups. Make sure this name doesn’t clash with the existing feature sets’ names.
- get_feature_names(): returns a list of feature names processed by the add-on.
- get_features(): receives a deep structure document (see Deep structure pos-tagged document) and returns a list of the features processed by the class. The features must be returned in the same order as the list given by get_feature_names().
The add-on needs to be placed in the PYTHON_PATH so that lingtools is able to find and import it. Alternatively, it may be available in another package installed in the system.
Add-ons can be included in the FeatureExtractor object using the following dictionary format:
addon_extra = {
'name': 'Extra', # Human-readable name for the feature set
'code': 'extra', # Code for the feature set
'package': None, # Package where the addon comes from, if any
'module': 'extra', # Name of the module
'class': 'ExtraFeatures', # Class providing the add-on
'options': {} # Parameters to be passed to the class, if any
}
And the feature extractor object can be initialized with one or more add-ons:
# Create a dictionary with settings, including the tokenizer object created previously
settings_with_addon = { "feature_sets": ["anew", "biber", "mrc"], # Choose which modules to extract
"vectorspace_name": 'glove', # Non-default vector space (optional)
"tokenizer_obj": tokenizer_obj,
"addons": [addon_extra] }
# Instantiate feature extractor object
feature_extractor_2 = FeatureExtractor(**settings_with_addon)
Calculate cosine similarities¶
The class lingtools.cosinedists.CosineDists can be used to extract cosine similarity measurements between words/words, words/documents and documents/documents.
To generate a cosine similarity matrix between a list of words:
import pandas
from lingtools.cosinedists import CosineDists
# Create a dictionary with settings, including the tokenizer object created previously
cosdist_settings = {"vectorspace_name": 'glove', # Non-default vector space (optional)
"tokenizer_obj": tokenizer_obj
}
# Instantiate object
cosdist_obj = CosineDists(**cosdist_settings)
# Load a list of words
cosdist_obj.load_list([u"dog", u"cat", u"pizza"])
# Get matrix as a dictionary
results = cosdist_obj.get_matrix_as_dict()
# Use pandas to load results as a table
df = pandas.DataFrame.from_dict(results)
Nearest neighbors¶
The class lingtools.cosinedists.CosineDists can also be used to find nearest neighbors from a given word or document.
To extract nearest neighbors from a word or document:
from lingtools.cosinedists import CosineDists
# Create a dictionary with settings, including the tokenizer object created previously
cosdist_settings = {"vectorspace_name": 'glove', # Non-default vector space (optional)
"tokenizer_obj": tokenizer_obj
}
# Instantiate object
cosdist_obj = CosineDists(**cosdist_settings)
# Get nearest neighbors
result = cosdist_obj.get_nearest_neighbors(u"dog", n=10)
Dynamic themes¶
The class lingtools.dynamicthemes.DynamicThemes calculates cosine similarity values between a dynamically constructed theme and a list of documents.
To use the class:
from lingtools.dynamicthemes import DynamicThemes
# Create a dictionary with settings, including the tokenizer object created previously
dynthemes_settings = {"vectorspace_name": 'glove', # Non-default vector space (optional)
"tokenizer_obj": tokenizer_obj
}
# Instantiate object
dynthemes_obj = DynamicThemes(**dynthemes_settings)
# Load the theme
dynthemes_obj.load_theme("Coffee", ["coffee", "bean", "beverage", "hot"])
# Load the file
dynthemes_obj.load_file('test.csv',
sep=';',
encoding='utf-8',
skip=1, # Skip the header
file_type='csv')
# Get the results
for idx, r in dynthemes_obj.process():
print("Doc index: %s" % idx)
print("Result tuple: %s" % r)
Important
Note that theme words will NOT be preprocessed. So, for example, if the vector space was trained using uses lemmas (as is the case with the Default vector space), theme words need to be informed already as lemmas.
The results are tuples in the following format: (theme_word, cosine_similarity, normalized_count). The first item refers to all theme words combined, while the following items list the similarity and count for each theme word separately:
# output:
Doc index: 1
Result tuple: [ (u'ALL', 0.10212916346547951, 0.04452212307975347),
(['coffee'], 0.1343591231224161, 0.036151228037899),
(['bean'], 0.1163471933552616, 0.004967344310551007),
(['beverage'], 0.08517727367333683, 0.0018397571520559286),
(['hot'], 0.06555257055307234, 0.0015637935792475394)
]
Extract distinctive words¶
The class lingtools.freqanalyzer.FreqAnalyzer can be used to compare two groups of documents and extract the most distinctive words between them using the chi-square test. For details on this approach, see: https://de.dariah.eu/tatom/feature_selection.html.
One additional step that needs to be performed to use this class is calling lingtools.freqanalyzer.FreqAnalyzer.create_termdocmatrix() after both group files have been loaded and before the function lingtools.freqanalyzer.FreqAnalyzer.get_distinctive_words() is called. See the example:
from lingtools.freqanalyzer import FreqAnalyzer
# Create a dictionary with settings, including the tokenizer object created previously
freqanalyze_settings = {"tokenizer_obj": tokenizer_obj}
# Instantiate object
freqanalyze_obj = FreqAnalyzer(**freqanalyze_settings)
# Load file from the first group
freqanalyze_obj.load_file('group1.csv',
group_name="group1",
sep=';',
encoding='utf-8',
skip=1, # Skip the header
file_type='csv')
# Load file from the second group
freqanalyze_obj.load_file('group2.csv',
group_name="group2",
sep='\t',
encoding='utf-8', # Must be the same as the other file
skip=1, # Skip the header
file_type='csv')
# Generate term-document matrix
freqanalyze_obj.create_termdocmatrix()
# Get the result
result = freqanalyze_obj.get_distinctive_words('group1', 'group2')
Important
Both files must have the same encoding!
The function returns a list of tuples in the following format: (word, chi-square, pval, word rate (per 1000 words) in group 1, word rate (per 1000 words) in group 2, group where word appears most). If there are no distinctive words between the two groups (that is, the files are too similar), the list will be empty.
Example result:
# print result
# (word, chi-square, pval, word_rate_group1, word_rate_group2, from_group)
[ (u'use', 15.777249575551785, 7.125418302648217e-05, 0.032749304077288356, 0.6091370558375634, 2),
(u'good', 14.004340900039825, 0.00018238907279415775, 0.13099721630915342, 0.77834179357022, 2),
(u'product', 12.342331635540466, 0.0004428016171161547, 0.06549860815457671, 0.5752961082910322, 2),
(u'one', 6.531100082712987, 0.010600437725800424, 0.06549860815457671, 0.37225042301184436, 2),
(u'size', 5.940777502067824, 0.014794490914363133, 0.360242344850172, 0.0676818950930626, 1),
(u'look', 5.408346134152585, 0.020040694334323265, 0.5239888652366138, 0.1692047377326565, 1),
(u'pretty', 5.218757467144564, 0.022344511137499933, 0.26199443261830685, 0.0338409475465313, 1),
(u'small', 4.475395319418088, 0.0343862467947065, 0.4257409530047487, 0.1353637901861252, 1),
(u'long', 4.306586021505376, 0.03796507910969707, 0.22924512854101853, 0.0338409475465313, 1) ]