This pipeline processes single-cell RNA sequencing (scRNA-seq) data using the Piscem-Alevin-Fry workflow. It supports two execution modes:
- Serverless (AWS Lambda) — Parallel read mapping across multiple Lambda instances for large-scale speedup.
- On-server / Standalone (any Linux machine) — Reproduces the traditional on-server execution baseline from the paper, running the full pipeline locally with no cloud dependencies.
Both modes produce identical gene-by-cell count matrices.
Ling-Hong Hung, Niharika Nasam, Chris Biju, Wes Lloyd, Ka Yee Yeung. Single-cell RNA sequencing data processing using cloud-based serverless computing. Gigabyte 2026
The standalone script reproduces the "On-Server (SSD) Execution" baseline from the paper (Table 2). It runs the identical Piscem-Alevin-Fry pipeline on any Linux x86_64 machine:
git clone https://github.com/BioDepot/scRNA-serverless.git
cd scRNA-serverless
bash scripts/e2e_standalone_pbmc.sh pbmc1kEverything (tools, reference data, FASTQs) is downloaded automatically from public sources. The pre-built reference index is archived on Zenodo (DOI: 10.5281/zenodo.19375096). See the On-Server Pipeline Guide for details.
One script covers PBMC 1K, PBMC 10K, and the MSK KO dataset. After the one-time AWS setup in the Serverless Pipeline Guide:
bash scripts/e2e_serverless_pbmc.sh pbmc1k
bash scripts/e2e_serverless_pbmc.sh pbmc10k
bash scripts/e2e_serverless_pbmc.sh koThe driver launches an m5dn.8xlarge, stripes both NVMe disks as RAID 0, splits FASTQs with rapidgzip -P 8 and 2 lanes at a time, maps on Lambda, then runs alevin-fry on the instance.
| Guide | Description |
|---|---|
| On-Server Pipeline Guide | Run the on-server pipeline on any Linux machine — no credentials needed, everything downloaded automatically |
| Serverless Pipeline Guide | Step-by-step instructions to run the serverless pipeline on your own AWS account (requires AWS, us-east-2 region) |
| Minimal AMI and NVMe Setup | Build a 20 GiB seed AMI and safely configure EC2 instance-store RAID 0 |
| Asynchronous Lambda Runbook | Nonblocking shard submission, atomic S3 claims, duplicate delivery, recovery, and later RAD materialization |
| Reproducibility Notes | Automatic fallbacks for AWS account limits, configuration reference, and local disk requirements |
| Direct S3 RAD Materializer | Ranged-download implementation and PBMC 1K benchmark |
| Production Async Benchmark | Final PBMC and KO runtime and cost tables |
| KO Sample-Eager Benchmark | Four-way grouped materialization, Lambda profile, and retained evidence |
| Piscem Single-Shard NVMe Profile | Six-thread PBMC shard timings for index loading, mapping, and RAD output |
| Cloud Piscem Profiling Handoff | Reproduce and diagnose one PBMC 1K shard on cloud NVMe |