rsloop is a PyO3-based asyncio event loop implemented in Rust.
Each rsloop.Loop owns a dedicated Rust runtime thread for loop coordination
and I/O work. That thread runs an rsloop-specialized vibeio runtime, using
io_uring on Linux, IOCP on Windows, and native kqueue readiness on macOS.
Native-stream TCP reads and Unix-domain socket reads run on that runtime. On Unix, generic
TCP protocol readers use a second vibeio runtime on the Python loop thread
(io_uring on Linux), avoiding cross-thread delivery of each read. Non-TLS accepts
run on either runtime depending on where the server starts. Python callbacks,
tasks, and coroutines run on the thread that calls
run_forever() or run_until_complete() (usually the main Python thread).
The package exposes:
- a native extension module at
rsloop._loop - a Python wrapper in
python/rsloop/__init__.py rsloop.Loop,rsloop.EventLoopPolicy,rsloop.new_event_loop(),rsloop.run(...),rsloop.install(),rsloop.uninstall(), andrsloop.build_info()
Repository metadata currently targets Python >=3.10.
The native runtime requires Linux 6.1+, macOS 13+, or Windows 10+ so its hot
paths can rely on modern completion, timer, and scheduler primitives.
Free-threaded CPython (3.14t) is supported: the extension declares
gil_used = false, so importing it no longer re-enables the GIL. See
Free-Threaded CPython for what that does and does not
buy you.
Project documentation now lives in docs/.
If you are new to the repository, start with:
To browse the docs locally with MkDocs:
uvx --from mkdocs mkdocs serveFrom PyPI:
pip install rsloopWith uv:
uv add rsloopFrom conda-forge, using pixi:
pixi add rsloopSimple entry point:
import rsloop
async def main(): ...
rsloop.run(main())Install as the default asyncio event loop policy:
import asyncio
import rsloop
rsloop.install()
try:
asyncio.run(main())
finally:
rsloop.uninstall()Manual loop creation also works:
import asyncio
import rsloop
loop = rsloop.new_event_loop()
asyncio.set_event_loop(loop)
try:
loop.run_until_complete(...)
finally:
asyncio.set_event_loop(None)
loop.close()Importing rsloop also patches asyncio.set_event_loop() so Python 3.10 can
accept an rsloop.Loop instance, matching the behavior exercised by
tests/test_run.py.
rsloop now exposes a small Rust interop API for downstream PyO3 extensions.
That lets you write your own async Rust code, return it to Python as an
awaitable, and run it under the active rsloop event loop.
The public entry point is rsloop::rust_async:
get_current_locals(...)future_into_py(...)future_into_py_with_locals(...)local_future_into_py(...)local_future_into_py_with_locals(...)- re-exports of
TaskLocalsandinto_future_with_locals(...)
See examples/rust/README.md for a complete
extension example built with maturin.
The current codebase implements these user-facing areas.
Loop lifecycle and scheduling:
run_forever,run_until_complete,stop,closetime,is_running,is_closedget_debug,set_debugcall_soon,call_soon_threadsafe,call_later,call_at- returned
HandleandTimerHandleobjects withcancel()/cancelled()
Tasks, futures, and execution helpers:
create_future,create_taskset_task_factory,get_task_factoryset_exception_handler,get_exception_handler,call_exception_handler,default_exception_handlerset_default_executor,run_in_executorshutdown_asyncgens,shutdown_default_executor- callback execution under captured
contextvars.Context asyncio.get_running_loop()support while running onrslooprsloop.run(...)helper, withasyncio.run(..., loop_factory=...)integration on Python 3.12+
I/O and networking:
add_reader,remove_reader,add_writer,remove_writersock_recv,sock_recv_into,sock_sendall,sock_accept,sock_connectgetaddrinfo,getnameinfocreate_server,create_connectioncreate_unix_server,create_unix_connectionconnect_accepted_socket- returned
Serverobjects withclose(),is_serving(),get_loop(), andsockets() - returned
StreamTransportobjects withwrite(),writelines(),close(),abort(),is_closing(),write_eof(),can_write_eof(),get_extra_info(),get_protocol(),set_protocol(),pause_reading(),resume_reading(),is_reading()
Pipes, subprocesses, and signals:
connect_read_pipe,connect_write_pipesubprocess_exec,subprocess_shell- returned
ProcessTransportandProcessPipeTransportobjects - higher-level compatibility with
asyncio.create_subprocess_exec()andasyncio.create_subprocess_shell() - Unix subprocess options including
cwd,env,executable,pass_fds,start_new_session,process_group,user,group,extra_groups,umask, andrestore_signals add_signal_handler,remove_signal_handler
Profiling:
profile(...),profiler_running(),start_profiler(),stop_profiler()- opt-in transport counters through
transport_stats()andreset_transport_stats()
Set RSLOOP_TRANSPORT_STATS=1 before importing rsloop to enable the transport
counters. They report read completions and bytes, Python-thread read drains,
wakeups, staged and direct writes, and Windows completion-to-poll rebinds.
Counters remain disabled by default so diagnostics add only one predictable
branch to transport hot paths.
Importing rsloop patches asyncio.open_connection() and
asyncio.start_server() by default.
That import-time behavior is controlled by RSLOOP_USE_FAST_STREAMS and can be
disabled with:
export RSLOOP_USE_FAST_STREAMS=0The native fast-stream path is used only when:
- the running loop is an
rsloop.Loop sslis unset orNone
Otherwise rsloop falls back to the stdlib asyncio.streams helpers.
On that path the reader handed to your code is the native
PyFastStreamReader rather than asyncio.StreamReader. It implements the
reading surface protocols actually use:
read(n=-1),readexactly(n)readline(),readuntil(separator=b"\n"), including the tuple-of-separators form CPython 3.13+ acceptsat_eof(),exception(),feed_data(),feed_eof(),set_exception()
These match asyncio.StreamReader down to the exception types and their
attributes — IncompleteReadError.partial, LimitOverrunError.consumed, the
ValueError that readline() raises on limit overrun — and down to what is
left in the buffer afterwards.
tests/test_stream_reader.py pins that by
driving the same feed scripts through both readers and comparing the results.
The implementation lives in
src/transport/stream/fast.rs and
is backed by the lower level transport code in
src/transport/stream/mod.rs.
rsloop builds and runs on free-threaded CPython 3.14 (3.14t). The extension
declares #[pymodule(gil_used = false)], which is what keeps CPython from
silently switching the GIL back on for the whole process at import time:
import sys
import rsloop
assert not sys._is_gil_enabled()
assert rsloop.build_info()["free_threaded"]What that buys you is that separate rsloop.Loop instances on separate threads
run concurrently rather than taking turns. A loop is still single-threaded
internally, and asyncio objects are still not thread-safe, so the model is one
loop per thread — not one loop shared across threads. call_soon_threadsafe()
remains the supported way to hand work to a loop from another thread, and it
keeps its FIFO ordering guarantee.
The pieces that made this safe:
- the generic stream-reader fast path writes into
StreamReader._bufferthrough a raw pointer; the size read, resize, and copy now run inside a critical section on thatbytearray, so a concurrent mutation cannot leave the copy writing into a freed allocation - the ready-queue refill preserves scheduling order when a drain slice leaves older callbacks in the batch. Under the GIL a cross-thread producer could only enqueue while the loop thread was parked, so the reordering was essentially unreachable; without the GIL producers append throughout the drain and it became routine
tests/test_free_threading.py covers this: parallel loops over both the native
and stdlib stream reader paths, call_soon_threadsafe() fan-in from eight
threads, and a check that importing rsloop leaves the GIL off.
Wheels are built for 3.14t alongside the GIL builds, and the test matrix runs
it as its own entry.
Each loop combines a coordination runtime with a loop-thread I/O runtime:
- the coordination thread handles commands, timers, and cross-thread work
- on Unix, generic TCP protocol readers run on the Python loop thread; native fast streams and Unix-domain readers retain coordination-thread I/O
- non-TLS accept loops use
vibeioon the thread that starts them - bounded ready-callback turns service loop-thread I/O even when Python tasks
continually yield with
sleep(0) - Windows TCP transports, including custom
asyncio.Protocolimplementations, start in IOCP completion mode and rebind to readiness mode beforestart_tlssynchronously reclaims a socket - generic
add_reader/add_writerdescriptors use cancellable OS-poll workers becausevibeiodoes not expose arbitrary raw-descriptor registration - some transport paths still fall back to helper threads, especially TLS I/O, TLS server accept, and parts of the legacy transport write path
The runtime dependency is now unified, but the codebase has not finished eliminating every helper thread yet.
Transport overload safeguards use conservative defaults: inbound reads pause
at 1 MiB of pending data per connection, buffered writes are capped at 64 MiB,
and a TLS server admits at most 256 simultaneous handshakes. The last two limits
can be adjusted before importing rsloop with
RSLOOP_MAX_WRITE_BUFFER_BYTES and RSLOOP_MAX_PENDING_TLS_HANDSHAKES.
These gaps are visible in the current implementation.
- TLS uses a
rustlsbackend with a narrower compatibility surface than CPython's OpenSSL-backedsslmodule. In particular, encrypted private keys are not supported yet, and the fast-stream monkeypatch still falls back to stdlib helpers wheneversslis enabled. TLS transport internals also still use helper-thread paths instead of the runtime-threadvibeiosocket path. - Subprocess support still has one notable gap:
preexec_fnremains unsupported because running arbitrary Python betweenfork()andexec()is unsafe in this runtime model. - Unix-specific APIs remain Unix-specific:
create_unix_server,create_unix_connection,add_signal_handler,remove_signal_handler. - Platform-specific limitations still apply:
Unix socket APIs and Unix signal handlers remain Unix-only, and several
subprocess options such as
pass_fds,user,group, andumaskare still specific to Unix process spawning. - The transport runtime model is still in transition: protocol readers on Unix avoid a coordination-thread hop, but native streams, generic descriptor watches, and TLS-heavy paths do not share one single-threaded I/O path.
Quick check:
cargo checkRelease build and editable install:
cargo build --release
uv run --with maturin maturin develop --releaseBuild release wheels into dist/wheels:
scripts/build-wheels.shOptionally build wheels with profile-guided optimization:
rustup component add llvm-tools-preview
scripts/build-pgo-wheels.shFor each requested Python ABI, the PGO wrapper creates an instrumented wheel, trains it on sustained HTTP, TLS, WebSocket, mixed-stream, bulk-transfer, idle-connection, callback, task, and TCP workloads, merges the resulting LLVM profiles, and builds that ABI's final wheel with its matching profile. Per-ABI training avoids discarding counters when PyO3's generated control flow differs between Python versions or free-threaded builds. The target must be native because the instrumented extension runs during training.
Set RSLOOP_PGO_SCENARIOS to override the comma-separated network scenarios.
The Wheels CI workflow disables PGO by default: tagged releases and ordinary
manual runs use the normal release-wheel builder. To opt in, enable the pgo
checkbox when manually running the workflow. LLVM tools are installed only for
PGO runs; source-distribution and publishing steps are unchanged.
When enabled, PGO is used on every supported platform except Windows ARM64.
Rust profile-generation binaries currently
crash on that target (rust-lang/rust#156675),
so it temporarily falls back to the normal fat-LTO release build.
scripts/build-wheels.sh currently defaults to
CPython 3.10 3.11 3.12 3.13 3.14, and
uses uv python install / uv python find to locate interpreters.
Profiling is behind the Cargo feature profiler and is disabled by default.
Build or install with that feature first:
cargo build --release --features profiler
uv run --with maturin maturin develop --release --features profilerThen wrap the code you want to inspect:
import rsloop
with rsloop.profile():
rsloop.run(main())Or manage the session manually:
import rsloop
rsloop.start_profiler()
try:
rsloop.run(main())
finally:
rsloop.stop_profiler()This starts a Tracy client inside the process. Build a release binary, open the Tracy desktop profiler, then connect to the running process while the profiled code is executing.
Release wheels do not include profiler support. Build locally with
--features profiler to enable it. The Tracy feature set is aimed at local
profiling: enable, only-localhost, and sampling.
For very short-lived runs you can force the process to block on exit until a
server has connected and drained all data by setting TRACY_NO_EXIT=1 in the
environment.
If the extension was built without --features profiler, profile() and
start_profiler() raise a runtime error.
Run the repository examples from the project root:
uv run python examples/01_basics.py
uv run python examples/02_fd_and_sockets.py
uv run python examples/03_streams.py
uv run python examples/04_unix_and_accepted_socket.py
uv run python examples/05_pipes_signals_subprocesses.pyExample files:
examples/01_basics.py,
examples/02_fd_and_sockets.py,
examples/03_streams.py,
examples/04_unix_and_accepted_socket.py,
examples/05_pipes_signals_subprocesses.py.
The repository also includes:
examples/fastapi_service.pyfor running the same FastAPI app on stdlibasyncio,uvloop, orrsloopbenches/compare_event_loops.pyfor callback, task, and TCP stream comparisons
uv run --with maturin maturin develop --release
uv run --with uvloop python benches/compare_event_loops.pyAn example output from that script on macOS (arm64) with CPython 3.14:
callbacks (200,000 ops)
loop median_s best_s ops_per_s peak_rss vs_fastest slower_by
rsloop 0.033083 0.032710 6,045,401 67.5 MiB 1.00x 0.0%
uvloop 0.040958 0.040721 4,883,026 72.8 MiB 1.24x 23.8%
asyncio 0.082233 0.082093 2,432,114 65.3 MiB 2.49x 148.6%
tasks (50,000 ops)
loop median_s best_s ops_per_s peak_rss vs_fastest slower_by
rsloop 0.063593 0.063286 786,247 37.6 MiB 1.00x 0.0%
uvloop 0.069614 0.069420 718,251 38.4 MiB 1.09x 9.5%
asyncio 0.108114 0.107502 462,473 36.1 MiB 1.70x 70.0%
tcp_streams (5,000 ops)
loop median_s best_s ops_per_s peak_rss vs_fastest slower_by
rsloop 0.090940 0.083355 54,981 32.2 MiB 1.00x 0.0%
uvloop 0.133182 0.127404 37,543 31.5 MiB 1.46x 46.5%
asyncio 0.302337 0.299813 16,538 29.6 MiB 3.32x 232.5%
The production-shaped workload matrix exercises HTTP, WebSocket libraries, TLS, mixed message sizes, backpressure, and connection lifecycle behavior:
uv run --with uvloop python benches/workload_matrix.py \
--loops rsloop,uvloop \
--sustained \
--json-output target/matrix-opt-final.jsonMeasured on September 6, 2026 with an Intel Core i9-9900K, Linux
7.0.0-31-generic (x86_64), CPython 3.14.0, rsloop 0.1.47 (release build), and
uvloop 0.22.1. Each row reports the median of seven measured runs after two
warmups, using 16 concurrent connections and 500 requests per
connection. Throughput is traffic-only operations per second, except for
bulk_transfer, which reports traffic MiB/s. The p95 columns are the medians
of each run's p95 latency; the difference is (rsloop / uvloop - 1) × 100%.
WebSocket library versions were websockets 17.0.1, aiohttp 3.14.3, Starlette 1.6.0, and uvicorn 0.52.3. The run used unrestricted CPU affinity, with other host services running but no concurrent builds or tests. These measurements are from a different host than the macOS microbenchmark example above.
| Scenario | rsloop | uvloop | rsloop difference | rsloop p95 | uvloop p95 |
|---|---|---|---|---|---|
| HTTP keep-alive | 51,561 | 51,410 | +0.3% | 0.354 ms | 0.351 ms |
| TLS HTTP | 72,465 | 25,493 | +184.3% | 0.241 ms | 0.668 ms |
| Raw WebSocket | 5,197 | 5,296 | -1.9% | 5.023 ms | 3.504 ms |
| Raw WebSocket over TLS | 5,395 | 4,824 | +11.8% | 3.945 ms | 3.870 ms |
websockets |
22,756 | 24,780 | -8.2% | 0.830 ms | 0.690 ms |
websockets over TLS |
26,044 | 15,081 | +72.7% | 0.672 ms | 1.109 ms |
| aiohttp WebSocket | 29,879 | 33,895 | -11.8% | 0.656 ms | 0.511 ms |
| aiohttp WebSocket over TLS | 32,983 | 19,063 | +73.0% | 0.526 ms | 0.890 ms |
| Starlette WebSocket | 18,524 | 20,325 | -8.9% | 0.997 ms | 0.832 ms |
| Starlette WebSocket over TLS | 18,058 | 13,555 | +33.2% | 0.942 ms | 1.257 ms |
| Mixed streams | 42,797 | 34,642 | +23.5% | 0.484 ms | 0.521 ms |
| Bulk transfer (MiB/s) | 1,993.5 | 1,223.7 | +62.9% | 14.768 ms | 26.061 ms |
The former single-burst idle-activation row has been retired: its traffic phase lasted only a few milliseconds and produced unstable throughput rankings. Idle activation now has a separate, versioned latency benchmark described below. The other measurements above are unchanged. HTTP's 0.3% difference is too small to call a win.
Compared with the same sustained workload on the pre-optimization build,
plain-text websockets, aiohttp, and Starlette throughput improved by 10.2%,
19.4%, and 22.3%, respectively. They still trail uvloop. The historical regression
gate also flagged HTTP tail latency and legacy idle activation; this is not an
across-the-board performance win. See the
benchmark documentation for workload definitions and
reproduction commands.
The ordinary matrix defaults are intentionally short enough for local smoke
and CI runs. Even with --sustained, compare repeated runs before drawing
performance conclusions for a deployment — competing desktop load matters more
than it looks, because rsloop trades helper-thread CPU for loop-thread work and
so has more to lose when cores are contended.
.venv/bin/python benches/workload_matrix.py \
--loops rsloop,uvloop --scenarios idle_connections --repeat 8 \
--idle-cycles 100 --idle-warmup-cycles 5 --idle-seconds 0.2 \
--json-output target/idle-v2-paired.jsonIdle v2 reuses 200 established connections across repeated idle/wakeup cycles.
It measures all replies from one shared activation timestamp, including task
scheduling delay, and reports first/50%/95%/all-reply latency. Eight fresh-process
blocks alternate loop order; confidence intervals resample whole paired runs,
not individual connections. Results are classified as improved, regressed, or
inconclusive using a 5% practical threshold and an approximate 95% confidence
interval. The command takes about six minutes; use --idle-cycles 3 --idle-warmup-cycles 1 --idle-seconds 0.01 --repeat 1 for a smoke test only.
The new measurements cannot be compared with the retired ops/s row. See benchmark methodology and regression handling for timing definitions, host controls, raw distributions, and sample requirements.
Validation on the Linux/i9-9900K host above collected 800 measured cycles per loop in 16 distinct processes, with unrestricted affinity and no concurrent builds or test runs. These are medians across runs of each run's median cycle milestone, in milliseconds (lower is better):
| Loop | First reply | 50% replied | 95% replied | All replied |
|---|---|---|---|---|
| rsloop | 17.510 | 17.893 | 18.169 | 18.299 |
| uvloop | 18.072 | 18.584 | 19.029 | 19.076 |
The preselected paired comparison is inconclusive: the geometric mean run-level p95 latency change for rsloop versus uvloop is +9.6%, with a 95% bootstrap interval of [-4.8%, +32.2%]. This uses paired run ratios, not the ratio of the table's medians. Individual cycle-p95 latencies still form fast and slow clusters (rsloop 3.423–28.634 ms; uvloop 3.708–24.329 ms). The new benchmark exposes that uncertainty rather than declaring a throughput winner.
See benches/README.md for workload details and
extra flags, and examples/README.md for the FastAPI
loop comparison example.
rsloop builds on the Python asyncio model and is implemented with
PyO3 on the Rust side. Runtime and socket I/O are powered by
vibeio.
This project is licensed under the Apache License, Version 2.0. See
LICENSE for the full text.
