lingtools.filereader module

class lingtools.filereader.FileReader(logging_interval=None)[source]

Bases: lingtools.filereader.FileReaderBase

Base class for reading files or lists

Parameters:logging_interval (int) – output logging info every logging_interval documents
get_counter()[source]

Get the internal counter so that it is possible to keep track of how many documents have been read.

get_docs()[source]

Generator function to iterate over the document source (whether list or file) while keeping track of the number of documents read so far via the internal counter.

Yields:(docid, document)
increment_counter(incr=1)[source]

Increment the internal counter by incr (default 1)

Parameters:incr (int) – increment step (default 1)
load_file(input_file, file_type=None, sep=None, encoding=None, skip=None, quotechar=None, paragraph_sep=None, clean_extra_spaces=False)[source]

Loads a file for processing.

In CSV files, it is expected that the first column contains the ID of the document, and the second column contains the text. All remaining columns are ignored.

Parameters:
  • input_file – path to the file.
  • type – txt_single|txt_multiple|csv. Default: txt_single
  • skip – How many lines to skip (in txt_multiple and csv files). Default: 0
  • paragraph_sep – New paragraph separator (in txt_single and zip files). Default:
  • sep – Separator in the CSV file. Default: ;
  • clean_extra_spaces – boolean. If true, double spaces and extra spaces between new lines will be removed. Default: False
  • encoding – Encoding of input file. Default: utf-8
  • quotechar – Quoting character (for csv). Default: b’”’
load_list(doc_list, doc_index=None, paragraph_sep=None, clean_extra_spaces=False)[source]

Loads a list of documents for processing. Note that if a file is loaded, the documents loaded by this function are ignored.

Parameters:
  • doc_list – a list of documents (unicode)
  • encoding – Encoding of input file. Default: utf-8
  • paragraph_sep – New paragraph separator (in txt_single and zip files). Default:
  • clean_extra_spaces – boolean. If true, double spaces and extra spaces between new lines will be removed. Default: False
reset_counter()[source]

Reset internal counter.

unload_file()[source]

Clear a previously loaded file (so that the object can be used with a list of documents instead)

class lingtools.filereader.FileReaderBase[source]

Bases: object

Base object for reading input files. Not to be used directly.

class lingtools.filereader.MultipleFileReader(logging_interval=100)[source]

Bases: lingtools.filereader.FileReaderBase

A class that is able to read from multiple sources, particularly to allow document comparison (as in lingtools.freqanalyzer.FreqAnalyzer).

get_docs()[source]

Generator function that iterates over all documents from all loaded document sources.

If there are multiple groups of documents, the group id will be prepended to the document id.

Yields:(docid, document)
load_file(input_file, group_name=None, file_type=None, paragraph_sep=None, sep=None, encoding=None, skip=None, quotechar=None, clean_extra_spaces=False)[source]

Loads a file for processing.

Parameters:
  • input_file – a CSV file containing an ID and the text itself. Remaining columns are ignored.
  • group_name – the name identifying this group of documents.
  • type – txt_single|txt_multiple|csv|zip. Default: txt_single
  • skip – How many lines to skip (in txt_multiple and csv files). Default: 0
  • paragraph_sep – New paragraph separator (in txt_single and zip files). Default:
  • sep – Separator in the CSV file. Default: ;
  • encoding – Encoding of input file. Default: utf-8
  • quotechar – Quoting character. Default: b’”’