- 2026-09-17: Open-sourced evaluation trajectories.
Accuracy (%) β. Baselines and StreamMind v1 are from the paper; π¦ marks new v2 runs (3,646 questions each, execution failures counted as zero). β means not evaluated or not supported.
| Setting | Method / backbone | RTP | HR | Tool | Pro |
|---|---|---|---|---|---|
| Offline | Qwen3.5-397B-A17B | 44.1 | 41.5 | 62.2 | β |
| Offline | MiMo-V2.5 | 38.0 | 35.8 | 47.9 | β |
| Offline | Kimi-K2.6 | 47.9 | 43.8 | 60.9 | β |
| Offline | Gemini 3.5 Flash | 51.3 | 51.4 | 70.8 | β |
| Offline | Qwen3.5-Omni | 41.8 | 35.8 | 49.4 | β |
| Recent-window | AURA / Qwen3-VL-8B-Instruct | 28.1 | 22.7 | β | 3.7 |
| Recent-window | MiniCPM-o-4.5 / Qwen3-8B | 22.1 | 9.8 | 17.1 | 7.5 |
| Text-summary | VST / Qwen2.5-VL-Instruct | 24.0 | 21.2 | β | β |
| Model-internal compression | StreamForest / Qwen2-7B | 17.9 | 14.4 | β | β |
| Model-internal compression | ThinkStream / Qwen2.5-VL-3B | 8.0 | 7.5 | 1.8 | 1.2 |
| Agentic streaming | StreamMind v1 / Qwen3.5-397B-A17B | 44.5 | 34.9 | 56.1 | 11.6 |
| π¦ Agentic streaming | StreamMind v2 β Qwen / Qwen3.5 | 44.9 | 40.8 | 57.1 | 19.4 |
| π¦ Agentic streaming | StreamMind v2 β Mixed / Gemini 3.5 Flash* | 55.9 | 49.1 | 70.2 | 23.3 |
StreamMind v2 (Mixed) largely matches the offline Gemini 3.5 Flash baseline, while operating online and supporting proactive interaction.
Mixed: Gemini 3.5 Flash for Front / Router / Search / Recall / Reviewer, Kimi-K2.6 for Memory Writer, and Qwen3.5 for Monitor.
StreamMind v2 adds OCR-indexed memory, keyword + embedding retrieval, and time-grounded Recall with visual verification while retaining the Front β Router β Recall / Search architecture. Scene-adaptive frame sampling and deduplication support continuous memory construction, while recalled ASR, on-screen text and historical frames provide evidence for answering.
Download trajectories on Hugging Face: baseline and StreamMind results, plus full v2 archives with Judge results, Debug traces, ASR, memory databases and evidence frames. Use the local v2 viewer to replay traces alongside original videos, without model API calls.
StreamArena_code/
βββ streamarena/ # shared library
β βββ data.py # dataset loader + EvalRecord/EvalOutput helpers
β βββ tools.py # Config, Frame, FrameBuffer, Serper ToolBox
β βββ media.py # frame extraction + subtitles + audio slicing + clip helpers
β βββ protocols.py # tool-call protocol parsers, thinking splitter
β βββ backends.py # OpenAIBackend + GeminiBackend
β βββ runner.py # unified multi-turn runner (offline backends)
β βββ judge.py # LLM-as-Judge scorer
βββ method/ # one folder per evaluated method
β βββ offline/ # Qwen / MiMo / Kimi / Qwen-Omni / Gemini (available)
β βββ streammind/ # Paper's StreamMind agent + streaming evaluation driver (available)
β βββ aura/ # AURA fixed-window streamer (client only; server = official vLLM)
β βββ minicpm/ # MiniCPM-o-4.5 unified (available; HF weights)
β βββ vst/ # VST text-summary streamer (available; HF weights)
β βββ streamforest/ # StreamForest native streamer (available; llava/ vendored)
β βββ thinkstream/ # ThinkStream native streamer (available; thinkstream/ vendored)
βββ judge/ # universal LLM-as-Judge scorer (available)
βββ requirements.txt
βββ LICENSE
βββ README.md
Every method writes the same records JSONL schema (see
method/offline/README.md for field details), so
the scorer in judge/ treats them all identically.
| Folder | Description | Status |
|---|---|---|
method/offline/ |
Turn-based MLLMs (Qwen / MiMo / Kimi / Qwen-Omni / Gemini) with the Serper agentic loop | Available |
method/streammind/ |
Paper's StreamMind agent with hierarchical memory, Recall / Search workers and proactive monitoring; streaming evaluation driver and StreamingAgent interface. |
Available |
method/aura/ |
AURA fixed-window streamer; client only, vLLM server per aurateam/AURA | Available |
method/minicpm/ |
MiniCPM-o-4.5 unified; weights openbmb/MiniCPM-o-4_5 | Available |
method/vst/ |
VST text-summary streamer; weights Catalan258/VST-7B | Available |
method/streamforest/ |
StreamForest native streamer; llava/ vendored from MCG-NJU/StreamForest (Apache-2.0) |
Available |
method/thinkstream/ |
ThinkStream native streamer; thinkstream/ vendored from CASIA-IVA-Lab/ThinkStream (MIT) |
Available |
judge/ |
LLM-as-Judge, shared across every method above | Available |
- 243 full-length videos, average duration 88.8 min, total ~300 GB.
- 3,646 timestamped open-ended questions across four capabilities:
Real-Time Perception (
RTP), Historical Retrospection (HR), External Tool Use (Tool), Proactive Interaction (Pro). - HR / Pro are stratified into horizon buckets up to >30 min for HR and >4 min for Pro.
- Two independent annotators + one blind auditor removed ~27% of drafts, leaving 3,646 validated tasks.
- StreamMind is the only streaming method evaluated on all four capabilities and reduces pooled query-to-answer latency by 66.2% on the same Qwen3.5-397B-A17B backbone.
Python 3.10+, FFmpeg, and enough disk space for the videos you want.
cd StreamArena_code
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtThe question set is a plain JSONL at the root of the dataset repo. Grab whichever language you need (English by default):
hf download hkuzxc/StreamArena \
--repo-type dataset \
--include "question.en.jsonl" \
--local-dir ./data/StreamArena
# for Chinese: --include "question.jsonl"Skip this step if you plan to pass the Hugging Face id directly
(DATASET=hkuzxc/StreamArena) -- the loader will fetch the JSONL for
you on first use.
One video (smoke test, ~1--2 GB):
hf download hkuzxc/StreamArena \
--repo-type dataset \
--include "videos/-J3qSQ2z4Nc.tar" \
--local-dir ./data/StreamArenaFull ~300 GB bundle:
hf download hkuzxc/StreamArena \
--repo-type dataset \
--include "videos/*.tar" \
--local-dir ./data/StreamArenaEach videos/<video_id>.tar bundles that video's .mp4 and any
subtitles. You have two options:
Option A -- do nothing. The runner auto-extracts a tar the first
time it needs the video (see _extract_tar_if_needed in
streamarena/data.py). Grabbing one tar is
enough for a smoke test.
Option B -- pre-extract everything. Handy when you want the full benchmark ready on a shared disk before starting a batch run:
cd ./data/StreamArena/videos
for f in *.tar; do
tar -xf "$f" && rm "$f" # drop `&& rm "$f"` if you want to keep the tars
done
cd -After extraction each video sits at
./data/StreamArena/videos/<video_id>/<video_id>.mp4 alongside its
.vtt subtitles, which is the layout the runner expects.
| Variable | Purpose |
|---|---|
GEMINI_API_URL / GEMINI_API_KEY |
Gemini endpoint + key (used by offline Gemini backend and the judge) |
SERPER_API_KEY |
Serper API key for text_search / image_search |
JUDGE_URL / JUDGE_API_KEY |
Judge endpoint + key |
STREAMARENA_DATASET |
Default dataset id / local path |
STREAMARENA_VIDEO_DIR |
Default video directory or tar archive folder |
STREAMARENA_LANGUAGE |
en or zh (default en) |
Two ways to test the Qwen backbones. The paper numbers come from option A; option B is provided for convenience when local GPUs are not available.
A) Local vLLM (the default; what the paper reports). Stand up Qwen3.5-397B-A17B behind an OpenAI-compatible endpoint following https://docs.vllm.ai/projects/ascend/en/latest/tutorials/models/Qwen3.5-397B-A17B.html, then point the driver at it:
export SERPER_API_KEY=...
export JUDGE_URL="https://generativelanguage.googleapis.com/v1beta/models/{model}:generateContent"
export JUDGE_API_KEY=...
AUTO_JUDGE=1 \
BASE_URL=http://localhost:12347/v1 MODEL=qwen3.5 \
VIDEO_DIR=./data/StreamArena/videos \
bash method/offline/scripts/run_qwen.shAll open-source backbones we report on (Qwen / MiMo / Kimi / MiniCPM / VST / StreamForest / ThinkStream / AURA) are served locally through vLLM or SGLang in the paper.
B) Qwen Official API (Alibaba Cloud, OpenAI-compatible mode). If you don't have GPUs, subscribe to a Qwen model at https://qwen.ai/apiplatform, then:
BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1 \
API_KEY="sk-..." \
MODEL=qwen3-max-preview \
VIDEO_DIR=./data/StreamArena/videos \
bash method/offline/scripts/run_qwen.shThe hosted API has its own model ids, rate limits, and request-schema quirks. Consult the official docs at https://qwen.ai/apiplatform and adjust the driver accordingly if something 4xx's -- we do not track hosted-API compatibility here.
Run + score separately:
bash method/offline/scripts/run_qwen.sh
bash judge/scripts/run_judge.sh method/offline/results/offline_qwen_qwen3.5.jsonl@misc{zhang2026streamarenacontinuousinteractivelonghorizon,
title={StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding},
author={Xichen Zhang and Guankai Li and Yinghao Zhu and Shijian Wang and Sitong Wu and Shaozuo Yu and Meng Chu and Yuan Lu and Jiaya Jia},
year={2026},
eprint={2608.05703},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.05703}
}Evaluation code is Apache-2.0. StreamArena annotations are CC-BY-NC-4.0. Video copyrights remain with the original YouTube uploaders; consult the dataset card before redistribution or commercial use.
Portions of two upstream baselines are vendored inside this repository
under method/*/_vendor/, together with their original LICENSE:
method/streamforest/_vendor/llava/is thellava/package from MCG-NJU/StreamForest (Apache-2.0). Seemethod/streamforest/_vendor/LICENSE.method/thinkstream/_vendor/thinkstream/is thethinkstream/package from CASIA-IVA-Lab/ThinkStream (MIT). Seemethod/thinkstream/_vendor/LICENSE.
AURA, MiniCPM-o-4.5, and VST are used through their published model checkpoints only; no code is vendored. See each folder's README for setup instructions and links to the upstream repos / weights.
