lingtools.featureextractor module

class lingtools.featureextractor.FeatureExtractor(features=None, feature_sets=None, addons=None, logging_interval=None, vectorspace_name=None, tokenizer_obj=None, config_file=None)[source]

Bases: object

Extracts linguistic features from text. It is possible to load an input file, or to give a single text document to process.

Parameters:
  • features – DEPRECATED
  • feature_sets (list) – list of features to be extracted from the text. Possible values: anew, basic, biber, extra, lcm, mrc, ner, semanticvectors, sentimentanalysis
  • addons (list<dict>) – a list of dictionary entries with the information about extra add-on modules to be loaded. See How to code with lingtools.
  • logging_interval (int) – output logging info every logging_interval documents
  • tokenizer_obj (LocalTokenizer) – initialized tokenizer object
  • vectorspace_name (string) – Name of the vector space to be used (see Configuration).
  • config_file (string) – Configuration file to use (optional. See Configuration).
get_feature_names()[source]

Returns the ordered, flat list of feature names that will be generated.

Returns:feature names
Return type:list
load_file(input_file, file_type=None, sep=None, encoding=None, skip=None, paragraph_sep=None, clean_extra_spaces=False)[source]

Loads a file for processing. See lingtools.filereader.FileReader.load_file() for a description of the options.

load_list(doc_list, doc_index=None, paragraph_sep=None, clean_extra_spaces=False)[source]

Loads a list of (unicode) documents for processing. See lingtools.filereader.FileReader.load_list() for a description of the options.

Note that the list of documents loaded here is ignored if a file has been loaded with load_file().

process()[source]

Processes previously loaded file or documents on the fly.

Yields a tuple (docid,featureContainer), which is loaded with a result container from the document.

Yields:a tuple (docid, featureContainer)
process_simple()[source]

Processes previously loaded file or documents on the fly.

The result is in the format of a tuple (docid, result_list), with result_list being a simple list with the results for each feature, in the order given by get_feature_names().

Yields:a tuple (docid, list_feature_values)
unload_file()[source]

Clear a previously loaded file. This function needs to be called if a file has been loaded before, but the user wants to process a list of documents (loaded with load_list()) instead.

class lingtools.featureextractor.FeaturesContainer(qualified_feature_sets)[source]

Bases: object

An object to contain extracted features from a text document.

Parameters:qualified_feature_sets – Dictionary of features to be contained
clear_results()[source]

Empty the results list in case it is needed to reuse the object.

get_results()[source]

Returns the structure with values. If called before “load_results()”, it will return the empty structure.

get_results_as_tuples()[source]

Returns the results as a list of tuples: (group_id, group_code, feature_id, feature_code, feature_value).

load_results(result_tuple)[source]

Receives a result tuple and returns a dictionary with feature values.

Parameters:result_tuple – a list of (feature_code, feature_value) tuples