SPRAS is a containerized library of pathway reconstruction tools. The framework contains different pathway reconstruction algorithms that connect proteins of interest in the context of a general protein-protein interaction network, allowing users to run multiple algorithms on their inputs. See the SPRAS manuscript or the GLBIO 2021 slides or video for more information.
This repository may undergo breaking changes. The instructions below support running SPRAS with a fixed configuration on example data to demonstrate its functionality. Open a GitHub issue or contact Anthony Gitter or Anna Ritz to provide feedback on this early version of SPRAS.
SPRAS is inspired by tools for single-cell transcriptomics such as BEELINE and dynverse that provide a unified interface to many related algorithms.
SPRAS overview showing PathLinker and Omics Integrator as representative pathway reconstruction algorithms.
We recommend running with at least 4 cores and 16 GB of RAM.
SPRAS runs on Linux, macOS, and Windows.
Our continuous integration runs on GitHub Actions using Ubuntu 24.04.5, macOS 26.6.2, Windows 10.0.26100.
The setup should take around 10 minutes.
SPRAS requires
- Files in this repository
- Python
- The Python packages listed in
environment.yml - Docker
First, download or clone this repository so that you have the Snakefile, example config file, and example data.
The easiest way to install Python and the required packages is with Anaconda.
The Carpentries Anaconda installation instructions provide guides and videos on how to install Anaconda for your operating system.
After installing Anaconda, you can run the following commands from the root directory of the spras repository
conda env create -f environment.yml
conda activate spras
to create a conda environment with the required packages and activate that environment. If you have a different version of Python already, you can install the specified versions of the required packages in your preferred manner instead of using Anaconda.
While the spras conda environment comes bundled with all of Python dependencies needed for spras to run, it does not yet have a working installation of spras itself.
To install spras in the environment, finish by running the following from the root directory of the repository:
python -m pip install .Use caution when pip installing directly to your computer without using some form of virtual/conda environment as this can alter your system's underlying Python modules, which could lead to unexpected behavior.
In most cases, you should only pip install spras if you're already working in the spras conda environment!
For developers, SPRAS can be installed via pip with the -e flag, as in python -m pip install -e .. This points Python back to the SPRAS repo so that any changes made to the source
code are reflected in the installed module.
You also need to install Docker. After installing Docker, start Docker before running SPRAS.
Running SPRAS locally on the test dataset takes around 10 minutes.
Once you have activated the conda environment and started Docker, you can run SPRAS on example data.
From the root directory of the spras repository, run the command
snakemake --cores 1 --configfile config/config.yaml
This will run the SPRAS workflow with the example config file (config/config.yaml) and input files.
The run covers 12 algorithms on two toy datasets, across parameter
combination listed in the config, followed by the post-processing analyses that
the config enables. Output files are written to the output directory.
You do not need to manually download Docker images from DockerHub before running SPRAS. The workflow will automatically download any missing images as long as Docker is running.
The example config defines two small datasets built from the files in input/:
data0:network.txtas the interactome, withnode-prizes.txt,sources.txt, andtargets.txtas node inputsdata1:alternative-network.txt, withnode-prizes.txt,sources.txt, andalternative-targets.txt
Four gold standard sets are declared (gs_nodes0.txt, gs_nodes1.txt,
gs_edges0.txt, gs_edges1.txt) and mapped to the two datasets.
Output is written to the output directory. For each dataset, expect:
- 19 pathway directories, one per algorithm-parameter combination, named
{dataset}-{algorithm}-params-{hash}. Across both datasets that is 38 reconstructions. {dataset}-pathway-summary.txt, node and edge counts and topological statistics for every output pathway for that dataset.{dataset}-ml/, the machine learning comparison of those output pathways.{dataset}-cytoscape.cys, a Cytoscape session holding every pathway graph for that dataset.{dataset}-{gold_standard}-eval/, one directory per dataset-gold standard pair declared ingold_standards, holding the corresponding evaluations.dataset-{dataset}-merged.pickle, the merged input data, plusgs-{gold_standard}-merged.picklefor each gold standard.logs/andprepared/, holding the parameter-hash mappings and the algorithm-specific input files generated before each run.
Large SPRAS workflows may benefit from execution with HTCondor, a scheduler/manager for distributed high-throughput computing workflows that allows many Snakemake steps to be run in parallel. For instructions on running SPRAS in this setting, see docker-wrappers/SPRAS/README.md.
Please refer to the instructions in our documentation.
Configuration file: Specifies which pathway reconstruction algorithms to run, which hyperparameter combinations to use, and which datasets to run them on.
Snakemake file: Defines a workflow to run all pathway reconstruction algorithms on all datasets with all specified hyperparameters.
Containerized pathway reconstruction algorithms: Pathway reconstruction algorithms are run via Docker images using the docker-py Python package.
The files to create these Docker images are in the docker-wrappers subdirectory along with links to algorithms' original repositories.
The Docker images are available on DockerHub.
Python wrapper for calling algorithms: Wrapper functions provide an interface between the common file formats for input and output data and the algorithm-specific file formats and reconstruction commands.
These wrappers are in the spras/ subdirectory.
Test code: Tests for the Docker wrappers and SPRAS code.
The tests require the conda environment in environment.yml and Docker.
Run the tests with pytest -s.
Some computing environments are unable to run Docker and prefer Singularity (or Apptainer) as the container runtime. SPRAS has limited experimental support for Singularity instead of Docker, and only for some pathway reconstruction algorithms. SPRAS uses the spython package to interface with Singularity, which only supports Linux.
SPRAS builds on public datasets and algorithms. If you use SPRAS in a research project, please cite the original datasets and algorithms in addition to SPRAS.
Part of ml.py is taken from the scikit-learn example code.
The original third-party code is available under the BSD 3-Clause License, Copyright © 2007 - 2023, scikit-learn developers.
A framework for benchmarking pathway reconstruction algorithms
Neha Talluri, Tristan Figueroa-Reid, Justin Hiemstra, Chris S Magnano, Adam Shedivy, Nistha Panda, Yancheng Liu, Sumedha Sanjeev, Oliver Faulkner Anderson, Altaf Barelvi, Aden O'Brien, Olivia T Johnson, James A Haddad, Spencer A Halberg-Spencer, Akniyet Nurbol, Iris Jan, Mahunan Degbelo, Daniel Nachreiner, Canek Llera-Magord, Gabriel Howland, Grace H Li, Anna Ritz+, Anthony Gitter+.
bioRxiv, 2026. doi:10.64898/2026.08.04.742550
+ denotes equal contribution.