Skip to contents

Creates a 2D track from tab-delimited file(s).

Usage

gtrack.2d.import(track = NULL, description = NULL, file = NULL)

Arguments

track

track name

description

a character string description

file

vector of file paths

Value

None.

Details

This function creates a 2D track track from one or more tab-delimited files. Each file must start with a header describing the columns. The first 6 columns must have the following names: 'chrom1', 'start1', 'end1', 'chrom2', 'start2', 'end2'. The last column is designated for the value and it may have an arbitrary name. The header is followed by a list of intervals and a value for each interval. Overlapping intervals are forbidden.

One can learn about the format of the tab-delimited file by running 'gextract' function on a 2D track with a 'file' parameter set to the name of the file.

If all the imported intervals represent a point (i.e. end == start + 1) a 'Points' track is created otherwise it is a 'Rectangles' track.

If 'gmultitasking' is on (the default), the input files are read by several processes in parallel, each process reading a contiguous subset of the files, and the chromosome pairs are then written in parallel, one pair per process at a time. A single input file is read by one process, so split a large input into several files to read it in parallel. The number of processes is bounded by 'gmax.processes'. The resulting track is identical to the one created with 'gmultitasking' off. A pair is started only when its estimated memory, added to that of the pairs being written, stays within 'gmax.mem.usage'. Larger pairs are taken first, but a smaller pair that fits may start while a larger one waits, and a pair estimated above 'gmax.mem.usage' is written alone. A pair with at least 4 x 'gmax.data.size' records is written in parts and estimated by its largest part, so lowering 'gmax.data.size' lets more pairs be written at once. Lower 'gmax.mem.usage' to use less memory. Reading the input is not covered by the estimate.

'description' is added as a track attribute.

Note: temporary files are created in the directory of the track during the run of the function. A few of them need to be kept simultaneously open. If the number of chromosomes and / or intervals is particularly high, a few thousands files might be needed to be opened simultaneously. Some operating systems limit the number of open files per user, in which case the function might fail with "Too many open files" or similar error. The workaround could be:

1. Increase the limit of simultaneously opened files (the way varies depending on your operating system). 2. Increase the value of 'gmax.data.size' option. Higher values of 'gmax.data.size' option will increased memory usage of the function but create fewer temporary files.