Quick start - glarue/intronIC GitHub Wiki
Quick start/testing
Installation
Using pip (recommended)
Install the last stable version from PyPI:
python -m pip install intronIC
Or install the latest version directly from GitHub:
python -m pip install git+https://github.com/glarue/intronIC
To upgrade to the latest version:
python -m pip install git+https://github.com/glarue/intronIC --upgrade
Using pixi (for development)
Pixi manages all dependencies automatically:
# Install pixi
curl -fsSL https://pixi.sh/install.sh | bash
# Clone and set up
git clone https://github.com/glarue/intronIC.git
cd intronIC
pixi install
pixi run intronIC --help
From source
Clone the repository and install in development mode:
git clone https://github.com/glarue/intronIC.git
cd intronIC
pip install -e .
Verifying Installation
After installing, verify it works with the bundled test data:
# Quick installation test (~1 minute with -p 4)
intronIC test -p 4
# Show where test data is located
intronIC test --show-only
This runs a smoke test to ensure intronIC is working correctly.
Dependencies
intronIC requires Python 3.10+ and the following packages:
- numpy
>=1.19.0: Numerical operations - scipy
>=1.5.0: Scientific computing - scikit-learn
>=0.22: SVM classifier - biogl
>=3.0: Bioinformatics utilities - matplotlib (optional): Plotting
- rich (optional): Progress bars
- pyyaml (optional): Configuration files
All required dependencies are installed automatically by pip.
intronIC was developed on Linux and has only been minimally tested on macOS and Windows.
Useful arguments
The required arguments for any classification run include a name (-n; see note below), along with:
- Genome (
-g) and annotation/BED (-a,-b) files or, - Intron sequences file (
-q) (see Training data and PWMS for formatting information, which matches the reference sequence format)
By default, intronIC includes non-canonical introns, considers only the longest isoform of each gene, and uses streaming mode for memory efficiency. Helpful arguments may include:
-
-pparallel processes, which reduce runtime -
-f cdsuse onlyCDSfeatures to identify introns (by default, uses bothCDSandexonfeatures) -
--no-ncexclude introns with non-canonical (non-GT-AG/GC-AG/AT-AC) boundaries -
-iinclude introns from multiple isoforms of the same gene (default: longest isoform only) -
--no-streamingdisable streaming mode (uses more memory but avoids temporary storage) -
--configpath to YAML configuration file for advanced settings
Configuration files
intronIC supports YAML configuration files for managing complex runs. Configuration files are searched in this order:
- Path specified by
--config .intronIC.yamlin current directory~/.config/intronIC/config.yaml~/.intronIC.yamlin home directory- Built-in defaults
CLI arguments always override config file values. To generate a template configuration file:
intronIC --generate-config > my_config.yaml
Example configuration:
scoring:
threshold: 90.0
exclude_noncanonical: false
extraction:
flank_length: 100
feature_type: both
performance:
processes: 8
Use with:
intronIC --config my_config.yaml -g genome.fa -a annotation.gff -n species
Running on test dataset
To test intronIC, use the bundled test data:
intronIC test -p 4
This automatically uses the included chromosome 19 test data and verifies your installation.
Manual test with custom data
If you prefer to manually test with specific files:
-
If you have installed via
pip, the test data is bundled with the package. UseintronIC test --show-onlyto see the location, or download the chromosome 19 FASTA and GFF3 sample files into a directory of your choice. -
If you have cloned the repo, first change to the
src/intronIC/data/test_datasubdirectory, which contains Ensembl annotations and sequence for chromosome 19 of the human genome. From the repo root, you can runpython -m intronICinstead ofintronICin the following examples.
Classify annotated introns
intronIC -g Homo_sapiens.Chr19.Ensembl_91.fa.gz -a Homo_sapiens.Chr19.Ensembl_91.gff3.gz -n homo_sapiens
The various output files contain different information about each intron; information can be cross-referenced by using the intron label (usually the first column of the file). An intron is classified U12-type when its type_id is u12 (equivalently, when its adjusted score is ≥ 50). The rel_score column (2nd column), centered on the 90% high-confidence threshold, marks the high-confidence U12-type subset when > 0 (see Output files for the full definition). For example, here is a U12-type AT-AC intron from the meta.iic file (16 tab-delimited columns: name, rel_score, dnts, motif_schematic, bp_context, bp_offset, length, parent, grandparent, index, family_size, frac_pos, phase, type_id, feature, attributes):
HomSap-ENSG00000141837@ENST00000614285_1(47);[c:-1] 10.0000 AT-AC GCC|ATATCCTTTT...TTTTCCTTAATT/TTTTCCTTAATT...AATAC|TCC CACCTCCAACACCCTTCTTTTCTTTGAACAAGAT[TTTTCCTTAATT]CCCAATAC -12 50719 ENST00000614285 ENSG00000141837 1 47 0.039 2 u12 cds corrected
To retrieve all U12-type introns from this file, filter on the type_id column (14th column), e.g.
awk -F'\t' '$14 == "u12"' homo_sapiens.meta.iic
To restrict to the high-confidence subset instead, filter on the rel_score column (rel_score > 0):
awk -F'\t' '$2 != "NA" && $2 > 0' homo_sapiens.meta.iic
Extract all annotated intron sequences
To retrieve all annotated intron sequences without classification, use the extract subcommand:
intronIC extract -g Homo_sapiens.Chr19.Ensembl_91.fa.gz -a Homo_sapiens.Chr19.Ensembl_91.gff3.gz -n homo_sapiens
See the rest of the Wiki for more extensive details about output files, usage info, etc.
A note on the -n (name) argument
By default, intronIC expects names in binomial (genus, species) form separated by a non-alphanumeric character, e.g. 'homo_sapiens', 'homo.sapiens', etc. intronIC then formats that name internally into a tag that it uses to label all output intron IDs, ignoring anything past the second non-alphanumeric character.
Output files, on the other hand, are named using the full name supplied via -n. If you'd prefer to have it leave whatever argument you supply to -n unmodified, use the --na flag.
If you are running multiple versions of the same species and would like to keep the same species abbreviations in the output intron data, add a tag to the end of the name, e.g. "homo_sapiens.v2"; the tags within files will be consistent ("HomSap"), but the file names across runs will be distinct.
Resource usage
Streaming vs in-memory
--streaming (the default) writes intron sequences to temporary on-disk storage during extraction and keeps only scoring motifs in memory; --in-memory keeps everything in memory. The two modes produce bit-identical classifications (locked in by integration tests as of v2.4); they differ only in the runtime/memory tradeoff.
To disable streaming (uses more memory but avoids temporary storage):
intronIC -g genome.fa -a annotation.gff -n species --in-memory
Benchmark
For reference wall-clock and peak-memory figures (Drosophila and full human), see Technical details — Memory and performance.
Scaling
Memory and runtime scale with the number of annotated introns rather than genome size. For non-human genomes:
- Non-model genomes (typical): ~1-5 GB peak, ~3-10 min (
-p 8) - Small test datasets: well under 1 GB, ~1-3 min
These estimates are for classification with the default v3 bundle (the single 42-model pmotif_adjudicated RBF SVM ensemble). Model training with intronIC train can take significantly longer (minutes to hours) depending on configuration. Using more parallel processes (-p N) reduces runtime in the scoring phase but extraction is largely I/O-bound.