Configuration

Lingtools will look for a configuration file at ~/.lingtools/config.yml by default. A sample configuration file is placed at ~/.lingtools/config.yml.sample. You can rename and edit it as needed.

Alternatively, it is possible to inform an alternative configuration file path to most lingtools classes. Check in the lingtools API which classes accept this option.

File format

The format of the configuration file is YAML (YAML Ain’t Markup Language), which is a human-readable data serialization language. In YAML, key value pairs within a map are separated by a colon, and structure is shown through indentation (one or more spaces). Comments begin with the number sign (#), can start anywhere on a line and continue until the end of the line. More information about YAML can be found at the official website: http://yaml.org/start.html.

The lingtools configuration file accepts the following settings, divided in three groups:

  • memory
    • limit
  • assets
    • anew
    • mrc
    • lcm
  • vector_spaces
    • [vector space entries]

Memory

Check the maximum memory allocation for the module (in bytes). The default value is 2 GiB, which might not be enough to analyze large files.

memory:
    limit: 2147483648

Assets

Assets location

You can configure the path in which lingtools assets can be found. By default, they will be placed in the folder .lingtools, in the user’s home folder.

assets:
    path: /home/USER/.lingtools

Assets

Lingtools requires some assets to run. These should be downloaded (see installation instructions) and placed in the configured path (see above).

The ANEW, MRC and LCM dictionaries are expected to be provided as a marisa trie (http://marisa-trie.readthedocs.io) pickled object.

ANEW dictionary

The ANEW norms of English Language. Configuration key: anew.

MRC dictionary

The MRC Psycholinguistic Database _http://websites.psychology.uwa.edu.au/school/MRCDatabase/mrc2.html. Configuration key mrc.

LCM dictionary

Linguistic category Model dictionary. Configuration key lcm.

Vector spaces configuration

In the configuration file, you can inform vector spaces to be made available to lingtools. It is possible to inform as many vector spaces as desired. For more information about vector spaces and the file formats that can be used, see Vector spaces in the page How to code with lingtools.

Each vector space entry (marked with indentation and an hyphen) must contain the following fields:

  • name: the vector space name (without spaces)
  • description: a one-line description of the vector space
  • dict: dictionary for the vector space, which can be a text file with one word per line or a gensim dictionary (https://radimrehurek.com/gensim/).
  • vectors: vector space file, a dense matrix of N-dimensional vectors, in which each row represents a word. There must be a 1:1 match between the vectors in this file and the words in the dictionary file, in correct order.

Hint

Loading gensim dictionaries is significantly faster than loading dictionaries from plain text. If speed is a concern, convert the dictionary file to a gensim dictionary. Important: ensure that the indices of the gensim dictionary match the vector space exactly!

There are many pre-trained vector spaces available online that can be used with lingtools, in English and in other languages. For example:

In most cases, the pre-trained vector spaces are available in the format of rows of a dense matrix stored in tab-delimited format (first element of each line corresponds to a word, followed by the values in the vector representing it). This format has to be processed as to extract the first column (the words) into a separate file, leaving the dense matrix on its own to be imported directly by numpy.

An example of configuration for two vector spaces in YAML format:

vector_spaces:
    - name: fasttext
      description: Fast Text, trained on Wikipedia English, 300 dimensions
      dict: PATH_TO_DICT_FILE
      vectors: PATH_TO_VECTOR_FILE
    - name: glove
      description: Global Vectors, 6B tokens, uncased, 300 dimensions
      dict: PATH_TO_DICT_FILE
      vectors: PATH_TO_VECTOR_FILE

These vector spaces are made available for the modules in addition to the default LSA space (see Default vector space).

Default vector space

The default vector space is an 300 dimension LSA model trained using the TASA corpus, lemmas only. For details, see the description of the model TASA 1 in the paper: Ştefănescu, D., Banjade, R., Rus, V.: Latent Semantic Analysis Models on Wikipedia and TASA, LREC (2014), available at http://deeptutor2.memphis.edu/Semilar-Web/public/lsa-models-lrec2014.html.

Example configuration file

memory:
    limit: 2147483648 # 2 * 1024 * 1024 * 1024 = 2 GiB

assets:
    path: /home/USER/.lingtools
    anew: /home/USER/.lingtools/assets/anew.marisa.pickle
    mrc: /home/USER/.lingtools/assets/mrc2.marisa.pickle
    lcm: /home/USER/.lingtools/assets/lcm.marisa.pickle

vector_spaces:
    - name: fasttext
      description: Fast Text, trained on Wikipedia English, 300 dimensions
      dict: /home/USER/.lingtools/assets/fasttext_dict
      vectors: /home/USER/.lingtools/assets/fasttext_model
    - name: glove
      description: Global Vectors, 6B tokens, uncased, 300 dimensions
      dict: /home/USER/.lingtools/assets/glove6B_dict
      vectors: /home/USER/.lingtools/assets/glove6B_model