Quick start¶
Installation¶
From PyPI¶
From source (for development)¶
vepyr requires a Rust toolchain and Python 3.10+.
curl -LsSf https://astral.sh/uv/install.sh | sh
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
- Clone and build:
git clone git@github.com:biodatageeks/vepyr.git
cd vepyr
RUSTFLAGS="-C target-cpu=native" uv sync --reinstall-package vepyr
- Verify:
Getting a cache¶
Annotation needs an Ensembl VEP cache in vepyr's optimized Parquet format. You have two options: download a prebuilt one (minutes, recommended) or build your own from an Ensembl VEP offline cache (hours of CPU).
Option A — download a prebuilt cache (recommended)¶
Release-116 GRCh38 caches for all three transcript sets are published on Hugging Face. Install the client once:
Then pull the cache type you want:
The download is resumable and the client verifies integrity, so there is no separate checksum step. Budget 31–36 G of disk per cache.
Confirm the cache reports the release you expect before annotating — this opens only the named contig's shards, so it is fast:
import os
import vepyr
cache = os.path.expanduser("~/vepyr_cache/116_GRCh38_merged")
print(vepyr.cache_contig_identity(cache, "chr22", expected_cache_version="116"))
For a second mirror, per-contig partial downloads, and the four prebuilt plugin caches, see Download Ensembl VEP and plugin caches.
Option B — build your own cache¶
Use this for release 115, for a cache type or assembly that is not mirrored, or when you want to convert a VEP cache you already hold.
Download and convert automatically¶
import vepyr
results = vepyr.build_cache(
release=115,
cache_dir="/data/vepyr_cache",
cache_type="ensembl",
)
for path, rows in results:
print(f"{path}: {rows:,} rows")
This downloads the Ensembl VEP 115 cache for homo_sapiens / GRCh38 and converts it to a partitioned Parquet cache.
Convert a local cache¶
If you already have the Ensembl VEP cache unpacked locally:
results = vepyr.build_cache(
release=115,
cache_dir="/data/vepyr_cache",
cache_type="ensembl",
local_cache="/data/ensembl_vep/homo_sapiens/115_GRCh38",
)
To rebuild a single raw entity without converting the full cache:
results = vepyr.build_cache_entity(
release=116,
cache_dir="/data/vepyr_cache",
entity="motif",
cache_type="merged",
local_cache="/data/ensembl_vep/homo_sapiens_merged/116_GRCh38",
overwrite=True,
)
This uses the same strict release/source validation as build_cache().
Options¶
| Parameter | Default | Description |
|---|---|---|
partitions |
8 |
DataFusion partitions for parallel conversion |
species |
homo_sapiens |
Species name |
assembly |
GRCh38 |
Genome assembly |
cache_type |
required | Ensembl VEP cache type: ensembl, merged, or refseq |
Annotating variants¶
Basic annotation¶
import vepyr
lf = vepyr.annotate(
vcf="input.vcf.gz",
cache_dir="/data/vepyr_cache/parquet/115_GRCh38_ensembl",
check_existing=True,
af=True,
max_af=True,
)
df = lf.collect()
print(df.select("chrom", "start", "ref", "alt", "most_severe_consequence").head())
Full --everything mode¶
Enable all annotation features (80-field CSQ). Requires a reference FASTA:
lf = vepyr.annotate(
vcf="input.vcf.gz",
cache_dir="/data/vepyr_cache/parquet/115_GRCh38_ensembl",
everything=True,
reference_fasta="GRCh38.fa",
)
df = lf.collect()
print(f"{df.height} variants x {df.width} columns")
workers controls how many within-contig annotation pipelines run
concurrently. workers=1 is the serial path; workers > 1 requires a
tabix-indexed (bgzip + .tbi) input VCF.
df = vepyr.annotate(
"input.vcf.gz",
"/data/vepyr_cache/parquet/115_GRCh38_ensembl",
workers=4,
).collect()
Writing annotated VCF output¶
Write results directly to a VCF file instead of returning a LazyFrame: