Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

mia ecosystem roadmap

This is the public roadmap for the mia ecosystem — the Bioconductor-native toolkit for microbiome data science built on (Tree)SummarizedExperiment.

Live board: https://github.com/orgs/microbiome/projects/6

Why this roadmap exists

The mia ecosystem spans more than ten Bioconductor packages with 236,000 downloads in 2025, and mia downloads have grown more than tenfold in two years to over 10,000 per month. That growth has outpaced the assumptions the code was originally written under. This roadmap sets out how we intend to make the framework computationally scalable and AI-ready over the next 24 months, organised as four goals and sixteen deliverables across eight quarters.

The roadmap is public so that users can see what is coming, contributors can find work that matters, and funders and reviewers can see what we committed to and whether we delivered it.

How to read it

Each goal (G1–G4) has an outcome, four deliverables, and a success indicator measured at the end of the project. Each deliverable is a tracked issue in this repository, mirrored onto the board with its goal, quarter and lead. Day-to-day work in the package repositories is connected to the roadmap through the shared G1:–G4: labels, which exist in mia and OMA and are being extended across the ecosystem.

Goal Outcome in one line Lead
G1 Performance & scalability for population-scale microbiome data Population-scale analyses run on ordinary institutional hardware. Tuomas Borman
G2 AI-enhanced microbiome data science methods Expert microbiome workflows can be automated with custom AI tools. Tuomas Borman
G3 Open benchmarks for microbiome data science Open, reusable scalability benchmarks replace ad hoc comparisons. Thomaz F. S. Bastiaanssen
G4 Community adoption and software sustainability Users can rely on the framework and new contributors can sustain it. Leo Lahti

Timeline

Quarters are project quarters (Q1 = first quarter of the funded period), aligned to Bioconductor's biannual release cycle.

Deliverable Q1 Q2 Q3 Q4 Q5 Q6 Q7 Q8
D1.1 Prioritize scalability gaps across the mia ecosystem ██ ██
D1.2 Optimize operations for large microbiome abundance matrices ██ ██
D1.3 Parallel and on-disk computing for TreeSummarizedExperiment ██ ██
D1.4 Optional hardware-accelerated routines ██ ██
D2.1 Curated microbiome data science workflows for AI agent training ██ ██
D2.2 Documented workflows from public repositories into probabilistic programming ██ ██
D2.3 AI skills and capabilities for microbiome data science ██ ██
D2.4 iSEEtree extended with AI-enhanced tools ██ ██
D3.1 Extend MGnifyR to return tidyomics-compatible structures ██ ██
D3.2 Standardized performance metrics for core microbiome tasks ██ ██
D3.3 Reusable benchmark suite with versioned reference datasets ██ ██
D3.4 Baseline results for mia and comparator frameworks ██ ██
D4.1 Expand automated test coverage across miaverse packages ██ ██
D4.2 Update OMA textbook and vignettes for new features ██ ██ ██ ██
D4.3 Enhance interfacing with Python and Julia ██ ██ ██ ██
D4.4 Contributor-onboarding workshops and sprints ██ ██ ██ ██ ██ ██ ██ ██

G1: Performance & scalability for population-scale microbiome data

Outcome. Researchers can run population-scale microbiome analyses with tens of thousands of samples routinely on standard institutional hardware, without hitting the memory and runtime limits of current tools.

Lead. Tuomas Borman

D1.1 Prioritize scalability gaps across the mia ecosystem — Q1–Q2

Systematically profile the mia package ecosystem to identify where current implementations hit memory and runtime limits, and rank the gaps by user impact.

D1.2 Optimize operations for large microbiome abundance matrices — Q3–Q4

Re-implement the highest-impact operations identified in D1.1 so that population-scale abundance matrices can be handled without exhausting memory or runtime.

D1.3 Parallel and on-disk computing for TreeSummarizedExperiment — Q5–Q6

Add parallel execution and on-disk (out-of-memory) backends for TreeSummarizedExperiment objects so that datasets exceeding available RAM remain analysable.

D1.4 Optional hardware-accelerated routines — Q7–Q8

Add optional hardware acceleration (e.g. GPU) for the most computationally intensive operations identified in D1.2-D1.3, as an opt-in layer on top of the CPU implementations.

Success indicator (Q8). 90% reduction in runtime and peak memory on population-scale benchmark datasets, and at least one dataset of 100,000 samples analysed on standard institutional hardware.

G2: AI-enhanced microbiome data science methods

Outcome. Microbiome data scientists can scale up their expert work with custom AI tools for automated microbiome workflows.

Lead. Tuomas Borman

D2.1 Curated microbiome data science workflows for AI agent training — Q1–Q2

Compile a curated, openly licensed collection of end-to-end microbiome data science workflows suitable as training and grounding material for AI agents.

D2.2 Documented workflows from public repositories into probabilistic programming — Q3–Q4

Document the full path from public microbiome data repositories through mia and tidyomics into probabilistic programming workflows with Stan.

D2.3 AI skills and capabilities for microbiome data science — Q5–Q6

Train and publish AI skills for microbiome data science using the curated material from D2.1-D2.2.

D2.4 iSEEtree extended with AI-enhanced tools — Q7–Q8

Extend the interactive iSEEtree package with AI-enhanced tools based on D2.3, so that non-programmers can use AI-assisted microbiome analyses.

Success indicator (Q8). Collection of curated end-to-end workflows published, and 50% growth in iSEEtree usage.

G3: Open benchmarks for microbiome data science

Outcome. The microbiome data science community has open scalability benchmarks against which methods can be validated, replacing current non-standardized comparisons.

Lead. Thomaz F. S. Bastiaanssen

D3.1 Extend MGnifyR to return tidyomics-compatible structures — Q1–Q2

Extend MGnifyR so that data retrieved from MGnify can be returned directly in tidyomics-compatible structures, preparing open data resources for benchmarking.

D3.2 Standardized performance metrics for core microbiome tasks — Q3–Q4

Define a standardized set of performance metrics for core microbiome data science tasks, replacing the current ad hoc, non-standardized comparisons.

D3.3 Reusable benchmark suite with versioned reference datasets — Q5–Q6

Build and publish a reusable, open benchmark suite with versioned reference datasets that the community can run and extend.

D3.4 Baseline results for mia and comparator frameworks — Q7–Q8

Run and publish baseline benchmark results for mia and at least one comparator framework, seeding community adoption of the suite.

Success indicator (Q8). Benchmark suite publicly released, documented, and used in at least 2 external comparison or method-development studies.

G4: Community adoption and software sustainability

Outcome. New and existing users can find, learn, and rely on the framework, and new contributors can be onboarded to sustain it beyond the grant period.

Lead. Leo Lahti

D4.1 Expand automated test coverage across miaverse packages — Q1–Q2

Expand automated test coverage across the miaverse packages being upgraded, so that the performance work of Goal 1 cannot silently change results.

D4.2 Update OMA textbook and vignettes for new features — Q3–Q6

Update the OMA online textbook and package vignettes to cover the new scalability, benchmarking and AI features as they are released.

D4.3 Enhance interfacing with Python and Julia — Q5–Q8

Improve interoperability with other languages used for microbiome data science (Python, Julia) by advancing platform-agnostic data formats and programming standards.

Note

Quarter range is not specified in the work plan; scheduled Q5-Q8 on the roadmap.

D4.4 Contributor-onboarding workshops and sprints — Q1–Q8

Run contributor-onboarding workshops and code sprints throughout the project, building on the existing international workshop programme, to broaden the contributor base beyond the grant period.

Note

The work plan lists D4.4 as Q7-Q8, while the supplementary commits to at least four workshops and sprints per year. Shown here as continuous Q1-Q8 with a dedicated onboarding push in Q7-Q8.

Success indicator (Q8). Test coverage reaching 97%, continued growth in downloads of the mia framework, and 6 new active contributors onboarded to the project.

Scope boundary

This roadmap covers package-level engineering for the mia ecosystem and its interoperability packages (MGnifyR, iSEEtree/miaDash, OMA). It is complementary to, and does not duplicate, cross-tool consortium infrastructure work.

How to contribute

Sustainability

Performance improvements land in the standard codebase and are maintained under the Bioconductor release process. Curated AI-agent training material feeds Bioconductor's community-maintained ai-agent-skills. The benchmark suite is designed to be extended by the community, including at least one external comparator-framework maintainer. The OMA textbook is a community-maintained resource with roughly 30 contributors already, kept current through the training programme.

About

Public roadmap for the mia ecosystem of R/Bioconductor packages for microbiome data science

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors