SNPkit is an R package designed for manipulation, organization, and analysis of genotypic data, with a strong focus on integration with tools such as FImpute and PLINK.

It provides robust S4-based data structures for storing genotypes and marker maps, along with functions to combine different genotype panels, summarize data, and prepare files for imputation and selection pipelines.

Key capabilities:

  • Import Illumina FinalReport.txt files (any panel density) and merge multiple genotype panels into a single object.
  • Quality control on SNPs and samples (call rate, MAF, HWE, monomorphic, duplicated positions, chromosome filters).
  • Prepare and run FImpute imputation and export to PLINK.
  • PCA (runPCA()) and anticlustering (runAnticlusteringPCA()) utilities for exploring structure and building balanced groups (e.g. batch design).

📦 Installation

SNPkit depends on snpStats, which is distributed through Bioconductor. Install it first:

if (!requireNamespace("BiocManager", quietly = TRUE)) install.packages("BiocManager")
BiocManager::install("snpStats")

Then install the stable release from CRAN:

Or install the development version (latest features) from GitHub:

# install.packages("remotes")
remotes::install_github("viniciusjunqueira/SNPkit")

Optional: faster PCA

runPCA() and runAnticlusteringPCA() can use RSpectra for a much faster, low-memory truncated PCA on wide genotype data. It is optional — install it to enable the fast path:

install.packages("RSpectra")

📖 Documentation

The full package website with detailed function reference and vignettes is available at:

Key pages:


📄 License

SNPkit is licensed under the GPL-3 license.