lingtools.featureextractor module¶
-
class
lingtools.featureextractor.FeatureExtractor(features=None, feature_sets=None, addons=None, logging_interval=None, vectorspace_name=None, tokenizer_obj=None, config_file=None)[source]¶ Bases:
objectExtracts linguistic features from text. It is possible to load an input file, or to give a single text document to process.
Parameters: - features – DEPRECATED
- feature_sets (list) – list of features to be extracted from the text. Possible values: anew, basic, biber, extra, lcm, mrc, ner, semanticvectors, sentimentanalysis
- addons (list<dict>) – a list of dictionary entries with the information about extra add-on modules to be loaded. See How to code with lingtools.
- logging_interval (int) – output logging info every logging_interval documents
- tokenizer_obj (LocalTokenizer) – initialized tokenizer object
- vectorspace_name (string) – Name of the vector space to be used (see Configuration).
- config_file (string) – Configuration file to use (optional. See Configuration).
-
get_feature_names()[source]¶ Returns the ordered, flat list of feature names that will be generated.
Returns: feature names Return type: list
-
load_file(input_file, file_type=None, sep=None, encoding=None, skip=None, paragraph_sep=None, clean_extra_spaces=False)[source]¶ Loads a file for processing. See
lingtools.filereader.FileReader.load_file()for a description of the options.
-
load_list(doc_list, doc_index=None, paragraph_sep=None, clean_extra_spaces=False)[source]¶ Loads a list of (unicode) documents for processing. See
lingtools.filereader.FileReader.load_list()for a description of the options.Note that the list of documents loaded here is ignored if a file has been loaded with
load_file().
-
process()[source]¶ Processes previously loaded file or documents on the fly.
Yields a tuple (docid,featureContainer), which is loaded with a result container from the document.
Yields: a tuple (docid, featureContainer)
-
process_simple()[source]¶ Processes previously loaded file or documents on the fly.
The result is in the format of a tuple (docid, result_list), with result_list being a simple list with the results for each feature, in the order given by get_feature_names().
Yields: a tuple (docid, list_feature_values)
-
unload_file()[source]¶ Clear a previously loaded file. This function needs to be called if a file has been loaded before, but the user wants to process a list of documents (loaded with
load_list()) instead.
-
class
lingtools.featureextractor.FeaturesContainer(qualified_feature_sets)[source]¶ Bases:
objectAn object to contain extracted features from a text document.
Parameters: qualified_feature_sets – Dictionary of features to be contained -
get_results()[source]¶ Returns the structure with values. If called before “load_results()”, it will return the empty structure.
-