Skip to content

Latest commit

 

History

123 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Serverless scRNA Pipeline

This pipeline processes single-cell RNA sequencing (scRNA-seq) data using the Piscem-Alevin-Fry workflow. It supports two execution modes:

  1. Serverless (AWS Lambda) — Parallel read mapping across multiple Lambda instances for large-scale speedup.
  2. On-server / Standalone (any Linux machine) — Reproduces the traditional on-server execution baseline from the paper, running the full pipeline locally with no cloud dependencies.

Both modes produce identical gene-by-cell count matrices.


Citation

Ling-Hong Hung, Niharika Nasam, Chris Biju, Wes Lloyd, Ka Yee Yeung. Single-cell RNA sequencing data processing using cloud-based serverless computing. Gigabyte 2026

Quick start (on-server baseline)

The standalone script reproduces the "On-Server (SSD) Execution" baseline from the paper (Table 2). It runs the identical Piscem-Alevin-Fry pipeline on any Linux x86_64 machine:

git clone https://github.com/BioDepot/scRNA-serverless.git
cd scRNA-serverless
bash scripts/e2e_standalone_pbmc.sh pbmc1k

Everything (tools, reference data, FASTQs) is downloaded automatically from public sources. The pre-built reference index is archived on Zenodo (DOI: 10.5281/zenodo.19375096). See the On-Server Pipeline Guide for details.

Quick start (serverless)

One script covers PBMC 1K, PBMC 10K, and the MSK KO dataset. After the one-time AWS setup in the Serverless Pipeline Guide:

bash scripts/e2e_serverless_pbmc.sh pbmc1k
bash scripts/e2e_serverless_pbmc.sh pbmc10k
bash scripts/e2e_serverless_pbmc.sh ko

The driver launches an m5dn.8xlarge, stripes both NVMe disks as RAID 0, splits FASTQs with rapidgzip -P 8 and 2 lanes at a time, maps on Lambda, then runs alevin-fry on the instance.


Documentation

Guide Description
On-Server Pipeline Guide Run the on-server pipeline on any Linux machine — no credentials needed, everything downloaded automatically
Serverless Pipeline Guide Step-by-step instructions to run the serverless pipeline on your own AWS account (requires AWS, us-east-2 region)
Minimal AMI and NVMe Setup Build a 20 GiB seed AMI and safely configure EC2 instance-store RAID 0
Asynchronous Lambda Runbook Nonblocking shard submission, atomic S3 claims, duplicate delivery, recovery, and later RAD materialization
Reproducibility Notes Automatic fallbacks for AWS account limits, configuration reference, and local disk requirements
Direct S3 RAD Materializer Ranged-download implementation and PBMC 1K benchmark
Production Async Benchmark Final PBMC and KO runtime and cost tables
KO Sample-Eager Benchmark Four-way grouped materialization, Lambda profile, and retained evidence
Piscem Single-Shard NVMe Profile Six-thread PBMC shard timings for index loading, mapping, and RAD output
Cloud Piscem Profiling Handoff Reproduce and diagnose one PBMC 1K shard on cloud NVMe

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages