Skip to content

Latest commit

Β 

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

StreamArena Evaluation Toolkit

Project Page Dataset Trajectories Paper License: Apache 2.0

StreamArena overview

News

Leaderboard

Accuracy (%) ↑. Baselines and StreamMind v1 are from the paper; 🟦 marks new v2 runs (3,646 questions each, execution failures counted as zero). β€” means not evaluated or not supported.

Setting Method / backbone RTP HR Tool Pro
Offline Qwen3.5-397B-A17B 44.1 41.5 62.2 β€”
Offline MiMo-V2.5 38.0 35.8 47.9 β€”
Offline Kimi-K2.6 47.9 43.8 60.9 β€”
Offline Gemini 3.5 Flash 51.3 51.4 70.8 β€”
Offline Qwen3.5-Omni 41.8 35.8 49.4 β€”
Recent-window AURA / Qwen3-VL-8B-Instruct 28.1 22.7 β€” 3.7
Recent-window MiniCPM-o-4.5 / Qwen3-8B 22.1 9.8 17.1 7.5
Text-summary VST / Qwen2.5-VL-Instruct 24.0 21.2 β€” β€”
Model-internal compression StreamForest / Qwen2-7B 17.9 14.4 β€” β€”
Model-internal compression ThinkStream / Qwen2.5-VL-3B 8.0 7.5 1.8 1.2
Agentic streaming StreamMind v1 / Qwen3.5-397B-A17B 44.5 34.9 56.1 11.6
🟦 Agentic streaming StreamMind v2 β€” Qwen / Qwen3.5 44.9 40.8 57.1 19.4
🟦 Agentic streaming StreamMind v2 β€” Mixed / Gemini 3.5 Flash* 55.9 49.1 70.2 23.3

StreamMind v2 (Mixed) largely matches the offline Gemini 3.5 Flash baseline, while operating online and supporting proactive interaction.

Mixed: Gemini 3.5 Flash for Front / Router / Search / Recall / Reviewer, Kimi-K2.6 for Memory Writer, and Qwen3.5 for Monitor.

StreamMind v2

StreamMind v2 adds OCR-indexed memory, keyword + embedding retrieval, and time-grounded Recall with visual verification while retaining the Front β†’ Router β†’ Recall / Search architecture. Scene-adaptive frame sampling and deduplication support continuous memory construction, while recalled ASR, on-screen text and historical frames provide evidence for answering.

Trajectories and visualization

Download trajectories on Hugging Face: baseline and StreamMind results, plus full v2 archives with Judge results, Debug traces, ASR, memory databases and evidence frames. Use the local v2 viewer to replay traces alongside original videos, without model API calls.

Repository layout

StreamArena_code/
β”œβ”€β”€ streamarena/           # shared library
β”‚   β”œβ”€β”€ data.py            # dataset loader + EvalRecord/EvalOutput helpers
β”‚   β”œβ”€β”€ tools.py           # Config, Frame, FrameBuffer, Serper ToolBox
β”‚   β”œβ”€β”€ media.py           # frame extraction + subtitles + audio slicing + clip helpers
β”‚   β”œβ”€β”€ protocols.py       # tool-call protocol parsers, thinking splitter
β”‚   β”œβ”€β”€ backends.py        # OpenAIBackend + GeminiBackend
β”‚   β”œβ”€β”€ runner.py          # unified multi-turn runner (offline backends)
β”‚   └── judge.py           # LLM-as-Judge scorer
β”œβ”€β”€ method/                # one folder per evaluated method
β”‚   β”œβ”€β”€ offline/           # Qwen / MiMo / Kimi / Qwen-Omni / Gemini    (available)
β”‚   β”œβ”€β”€ streammind/        # Paper's StreamMind agent + streaming evaluation driver (available)
β”‚   β”œβ”€β”€ aura/              # AURA fixed-window streamer                 (client only; server = official vLLM)
β”‚   β”œβ”€β”€ minicpm/           # MiniCPM-o-4.5 unified                      (available; HF weights)
β”‚   β”œβ”€β”€ vst/               # VST text-summary streamer                  (available; HF weights)
β”‚   β”œβ”€β”€ streamforest/      # StreamForest native streamer               (available; llava/ vendored)
β”‚   └── thinkstream/       # ThinkStream native streamer                (available; thinkstream/ vendored)
β”œβ”€β”€ judge/                 # universal LLM-as-Judge scorer              (available)
β”œβ”€β”€ requirements.txt
β”œβ”€β”€ LICENSE
└── README.md

Every method writes the same records JSONL schema (see method/offline/README.md for field details), so the scorer in judge/ treats them all identically.

Methods status

Folder Description Status
method/offline/ Turn-based MLLMs (Qwen / MiMo / Kimi / Qwen-Omni / Gemini) with the Serper agentic loop Available
method/streammind/ Paper's StreamMind agent with hierarchical memory, Recall / Search workers and proactive monitoring; streaming evaluation driver and StreamingAgent interface. Available
method/aura/ AURA fixed-window streamer; client only, vLLM server per aurateam/AURA Available
method/minicpm/ MiniCPM-o-4.5 unified; weights openbmb/MiniCPM-o-4_5 Available
method/vst/ VST text-summary streamer; weights Catalan258/VST-7B Available
method/streamforest/ StreamForest native streamer; llava/ vendored from MCG-NJU/StreamForest (Apache-2.0) Available
method/thinkstream/ ThinkStream native streamer; thinkstream/ vendored from CASIA-IVA-Lab/ThinkStream (MIT) Available
judge/ LLM-as-Judge, shared across every method above Available

Paper highlights

  • 243 full-length videos, average duration 88.8 min, total ~300 GB.
  • 3,646 timestamped open-ended questions across four capabilities: Real-Time Perception (RTP), Historical Retrospection (HR), External Tool Use (Tool), Proactive Interaction (Pro).
  • HR / Pro are stratified into horizon buckets up to >30 min for HR and >4 min for Pro.
  • Two independent annotators + one blind auditor removed ~27% of drafts, leaving 3,646 validated tasks.
  • StreamMind is the only streaming method evaluated on all four capabilities and reduces pooled query-to-answer latency by 66.2% on the same Qwen3.5-397B-A17B backbone.

Installation

Python 3.10+, FFmpeg, and enough disk space for the videos you want.

cd StreamArena_code
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

1. Question annotations (always required)

The question set is a plain JSONL at the root of the dataset repo. Grab whichever language you need (English by default):

hf download hkuzxc/StreamArena \
  --repo-type dataset \
  --include "question.en.jsonl" \
  --local-dir ./data/StreamArena
# for Chinese: --include "question.jsonl"

Skip this step if you plan to pass the Hugging Face id directly (DATASET=hkuzxc/StreamArena) -- the loader will fetch the JSONL for you on first use.

2. Videos

One video (smoke test, ~1--2 GB):

hf download hkuzxc/StreamArena \
  --repo-type dataset \
  --include "videos/-J3qSQ2z4Nc.tar" \
  --local-dir ./data/StreamArena

Full ~300 GB bundle:

hf download hkuzxc/StreamArena \
  --repo-type dataset \
  --include "videos/*.tar" \
  --local-dir ./data/StreamArena

3. Extract video tars

Each videos/<video_id>.tar bundles that video's .mp4 and any subtitles. You have two options:

Option A -- do nothing. The runner auto-extracts a tar the first time it needs the video (see _extract_tar_if_needed in streamarena/data.py). Grabbing one tar is enough for a smoke test.

Option B -- pre-extract everything. Handy when you want the full benchmark ready on a shared disk before starting a batch run:

cd ./data/StreamArena/videos
for f in *.tar; do
  tar -xf "$f" && rm "$f"   # drop `&& rm "$f"` if you want to keep the tars
done
cd -

After extraction each video sits at ./data/StreamArena/videos/<video_id>/<video_id>.mp4 alongside its .vtt subtitles, which is the layout the runner expects.

Environment variables

Variable Purpose
GEMINI_API_URL / GEMINI_API_KEY Gemini endpoint + key (used by offline Gemini backend and the judge)
SERPER_API_KEY Serper API key for text_search / image_search
JUDGE_URL / JUDGE_API_KEY Judge endpoint + key
STREAMARENA_DATASET Default dataset id / local path
STREAMARENA_VIDEO_DIR Default video directory or tar archive folder
STREAMARENA_LANGUAGE en or zh (default en)

End-to-end example (Qwen + auto-judge)

Two ways to test the Qwen backbones. The paper numbers come from option A; option B is provided for convenience when local GPUs are not available.

A) Local vLLM (the default; what the paper reports). Stand up Qwen3.5-397B-A17B behind an OpenAI-compatible endpoint following https://docs.vllm.ai/projects/ascend/en/latest/tutorials/models/Qwen3.5-397B-A17B.html, then point the driver at it:

export SERPER_API_KEY=...
export JUDGE_URL="https://generativelanguage.googleapis.com/v1beta/models/{model}:generateContent"
export JUDGE_API_KEY=...

AUTO_JUDGE=1 \
BASE_URL=http://localhost:12347/v1 MODEL=qwen3.5 \
VIDEO_DIR=./data/StreamArena/videos \
bash method/offline/scripts/run_qwen.sh

All open-source backbones we report on (Qwen / MiMo / Kimi / MiniCPM / VST / StreamForest / ThinkStream / AURA) are served locally through vLLM or SGLang in the paper.

B) Qwen Official API (Alibaba Cloud, OpenAI-compatible mode). If you don't have GPUs, subscribe to a Qwen model at https://qwen.ai/apiplatform, then:

BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1 \
API_KEY="sk-..." \
MODEL=qwen3-max-preview \
VIDEO_DIR=./data/StreamArena/videos \
bash method/offline/scripts/run_qwen.sh

The hosted API has its own model ids, rate limits, and request-schema quirks. Consult the official docs at https://qwen.ai/apiplatform and adjust the driver accordingly if something 4xx's -- we do not track hosted-API compatibility here.

Run + score separately:

bash method/offline/scripts/run_qwen.sh
bash judge/scripts/run_judge.sh method/offline/results/offline_qwen_qwen3.5.jsonl

Citation

@misc{zhang2026streamarenacontinuousinteractivelonghorizon,
  title={StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding},
  author={Xichen Zhang and Guankai Li and Yinghao Zhu and Shijian Wang and Sitong Wu and Shaozuo Yu and Meng Chu and Yuan Lu and Jiaya Jia},
  year={2026},
  eprint={2608.05703},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2608.05703}
}

Licenses

Evaluation code is Apache-2.0. StreamArena annotations are CC-BY-NC-4.0. Video copyrights remain with the original YouTube uploaders; consult the dataset card before redistribution or commercial use.

Third-party notices

Portions of two upstream baselines are vendored inside this repository under method/*/_vendor/, together with their original LICENSE:

AURA, MiniCPM-o-4.5, and VST are used through their published model checkpoints only; no code is vendored. See each folder's README for setup instructions and links to the upstream repos / weights.

About

No description, website, or topics provided.

Resources

Stars

34 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages