Configuration¶
Lingtools will look for a configuration file at ~/.lingtools/config.yml by default. A sample configuration file is placed at ~/.lingtools/config.yml.sample. You can rename and edit it as needed.
Alternatively, it is possible to inform an alternative configuration file path to most lingtools classes. Check in the lingtools API which classes accept this option.
File format¶
The format of the configuration file is YAML (YAML Ain’t Markup Language), which is a human-readable data serialization language. In YAML, key value pairs within a map are separated by a colon, and structure is shown through indentation (one or more spaces). Comments begin with the number sign (#), can start anywhere on a line and continue until the end of the line. More information about YAML can be found at the official website: http://yaml.org/start.html.
The lingtools configuration file accepts the following settings, divided in three groups:
- memory
- limit
- assets
- anew
- mrc
- lcm
- vector_spaces
- [vector space entries]
Memory¶
Check the maximum memory allocation for the module (in bytes). The default value is 2 GiB, which might not be enough to analyze large files.
memory:
limit: 2147483648
Assets¶
Assets location¶
You can configure the path in which lingtools assets can be found. By default, they will be placed in the folder .lingtools, in the user’s home folder.
assets:
path: /home/USER/.lingtools
Assets¶
Lingtools requires some assets to run. These should be downloaded (see installation instructions) and placed in the configured path (see above).
The ANEW, MRC and LCM dictionaries are expected to be provided as a marisa trie (http://marisa-trie.readthedocs.io) pickled object.
ANEW dictionary¶
The ANEW norms of English Language. Configuration key: anew.
MRC dictionary¶
The MRC Psycholinguistic Database _http://websites.psychology.uwa.edu.au/school/MRCDatabase/mrc2.html. Configuration key mrc.
LCM dictionary¶
Linguistic category Model dictionary. Configuration key lcm.
Vector spaces configuration¶
In the configuration file, you can inform vector spaces to be made available to lingtools. It is possible to inform as many vector spaces as desired. For more information about vector spaces and the file formats that can be used, see Vector spaces in the page How to code with lingtools.
Each vector space entry (marked with indentation and an hyphen) must contain the following fields:
- name: the vector space name (without spaces)
- description: a one-line description of the vector space
- dict: dictionary for the vector space, which can be a text file with one word per line or a gensim dictionary (https://radimrehurek.com/gensim/).
- vectors: vector space file, a dense matrix of N-dimensional vectors, in which each row represents a word. There must be a 1:1 match between the vectors in this file and the words in the dictionary file, in correct order.
Hint
Loading gensim dictionaries is significantly faster than loading dictionaries from plain text. If speed is a concern, convert the dictionary file to a gensim dictionary. Important: ensure that the indices of the gensim dictionary match the vector space exactly!
There are many pre-trained vector spaces available online that can be used with lingtools, in English and in other languages. For example:
- GloVe: Global Vectors for Word Representation: https://nlp.stanford.edu/projects/glove/
- FastText: https://github.com/facebookresearch/fastText/blob/master/pretrained-vectors.md
- LSA spaces for R: http://www.lingexp.uni-tuebingen.de/z2/LSAspaces/ and https://sites.google.com/site/fritzgntr/software-resources
In most cases, the pre-trained vector spaces are available in the format of rows of a dense matrix stored in tab-delimited format (first element of each line corresponds to a word, followed by the values in the vector representing it). This format has to be processed as to extract the first column (the words) into a separate file, leaving the dense matrix on its own to be imported directly by numpy.
An example of configuration for two vector spaces in YAML format:
vector_spaces:
- name: fasttext
description: Fast Text, trained on Wikipedia English, 300 dimensions
dict: PATH_TO_DICT_FILE
vectors: PATH_TO_VECTOR_FILE
- name: glove
description: Global Vectors, 6B tokens, uncased, 300 dimensions
dict: PATH_TO_DICT_FILE
vectors: PATH_TO_VECTOR_FILE
These vector spaces are made available for the modules in addition to the default LSA space (see Default vector space).
Default vector space¶
The default vector space is an 300 dimension LSA model trained using the TASA corpus, lemmas only. For details, see the description of the model TASA 1 in the paper: Ştefănescu, D., Banjade, R., Rus, V.: Latent Semantic Analysis Models on Wikipedia and TASA, LREC (2014), available at http://deeptutor2.memphis.edu/Semilar-Web/public/lsa-models-lrec2014.html.
Example configuration file¶
memory:
limit: 2147483648 # 2 * 1024 * 1024 * 1024 = 2 GiB
assets:
path: /home/USER/.lingtools
anew: /home/USER/.lingtools/assets/anew.marisa.pickle
mrc: /home/USER/.lingtools/assets/mrc2.marisa.pickle
lcm: /home/USER/.lingtools/assets/lcm.marisa.pickle
vector_spaces:
- name: fasttext
description: Fast Text, trained on Wikipedia English, 300 dimensions
dict: /home/USER/.lingtools/assets/fasttext_dict
vectors: /home/USER/.lingtools/assets/fasttext_model
- name: glove
description: Global Vectors, 6B tokens, uncased, 300 dimensions
dict: /home/USER/.lingtools/assets/glove6B_dict
vectors: /home/USER/.lingtools/assets/glove6B_model