Skip to content

Fix six spatial-audio rendering defects - #164

Merged
ctoth merged 6 commits into
feat/media-occlusionfrom
fix/spatial-rendering
Oct 5, 2026
Merged

ctoth merged 6 commits into
feat/media-occlusionfrom
fix/spatial-rendering

Conversation

@ctoth

@ctoth ctoth commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Six spatial-audio rendering fixes, one commit each with its tests. Stacked on #163 (feat/media-occlusion), which touches the same file.

What was wrong, and what changes

1. A voice-chat speaker with no known position was placed at the room origin. LiveKitSpatialAudioBridge fell back to [0, 0, 0], so a phone-call partner, or a speaker still connected from the room just left, was heard as a point at the centre of the current room at floor height. A voice is now positional only while its speaker is an entity in the current scene. Anyone else is passed as null, and MediaVoices plays that voice as a stereo source at centre pan: it has no 3D panner, so no position, no distance model and no direction. When the entity appears or leaves, the voice's sound is rebuilt in the other mode on the same track, keeping its named route. The test that asserted origin placement is rewritten to the new rule.

2. A sound's position glided across a room change. MediaService glides sound positions on its own VectorTweener, a different instance from the Client.Spatial handler's, so the handler's cancel on Scene never reached it. Scene now calls MediaService.sceneChanged(): glides in flight land on their targets at once, and the first position each sound receives afterwards is placed directly.

3. A re-Play at a moved cursor destroyed and recreated the voice. The segment comparison included start, which is the join cursor, not part of the region. Region identity is now the repeat window, loopStart..finish. On a kept voice a start seeks only when the playhead is more than MEDIA_SEEK_TOLERANCE_MS (150 ms) from the cursor. continue: false still forces a new voice.

4. Panning mode could not change on a kept voice. The mode is fixed when Cacophony creates a source, and nothing handled a later change of is3d. A Play or an Update that changes is3d for a playing key now builds the same source and region in the new mode; the old voice plays on while it loads, and the new one starts at the old one's playhead with no second fade-in. Volume, gain, pitch, spatial profile, loop count, occlusion and routing carry over.

5. Two writers moved the engine listener. ListenerPosition glided it every frame at the previously stored velocity; the listener's own EntityMove wrote the target outright when its entity glide finished, at the new velocity. A mover gets both for one step, so the listener jumped to the target and back. Both now feed one glide on the listener tween key.

6. Position glides froze in a hidden tab. The tweener schedules with requestAnimationFrame, which browsers stop for hidden tabs while audio keeps playing. It now takes a visibility source (the document by default, injectable so it stays testable without a DOM): a tween requested while hidden snaps, and glides in flight land on their targets when the document becomes hidden.

Things a reviewer should know

  • Region identity (3) is not a bare field comparison. The MOO sends loopStart only for a sound that repeats. A single pass re-Played without it would never match if the client's defaulted loopStart (= start) were compared, so such a voice is kept when its region reaches back to the new cursor. A repeating sound re-Played without loopStart still gets a new voice: per the wire contract its repeat window follows start.
  • Media glides finish, they are not cancelled in place (2). Cancelling would leave a sound part-way along its glide, with no further position coming if the Scene was not a room change.
  • An Update's fields are now folded into the stored Play payload (4). A rebuild then replays the state in effect now, not the original Play's. This also applies to the existing moved-segment rebuild.
  • Not rebuilt on an is3d change: ambisonic sounds, and preloaded sounds that were never Played.
  • Segment repeat timers are not rescheduled on a kept re-Play (3). An existing test pins that behaviour.
  • Not verified in a browser. The engine is mocked in these tests. The voice rebuild in (1) recreates the stream source, and with it Cacophony's muted priming element; whether Chromium leaves an audible gap at that moment needs a listen.

Checks

  • npm run typecheck: clean
  • npm test: 128 files, 1516 tests passed
  • npm run test:audio-cache: 3 PASS lines

ctoth added 6 commits October 4, 2026 19:58
Browsers stop requestAnimationFrame for a hidden or minimised tab while
audio keeps playing, so a glide froze part-way until the tab was shown
again. The tweener now takes a visibility source (the document by
default, injectable so it stays testable without a DOM): a tween asked
for while hidden snaps to its target, and glides in flight when the
document becomes hidden land on their targets at once. Nothing is left
in flight, so nothing moves when the tab is shown again.

Adds VectorTweener.finishAll(), and an idle tweener now holds neither a
pending frame nor a visibility listener.
Sound positions glide on MediaService's own VectorTweener, a separate
instance from the Client.Spatial handler's, so the handler's cancel on
Client.Spatial.Scene never reached them: a sound kept gliding across a
room change, between two rooms' unrelated coordinates.

The Scene handler now calls MediaService.sceneChanged(). Glides in
flight land on their targets at once (the last position the server
gave, never somewhere part-way), and the first position each sound
receives afterwards is placed directly instead of gliding from a
position in the previous room's coordinates.
Two handlers wrote the Cacophony listener position. ListenerPosition
glided it every frame at the previously stored velocity; the listener's
own EntityMove wrote the target outright when its entity glide finished,
at the new velocity. The server sends the mover both for one step, so
with different speeds the listener jumped to the target and then back
onto the slower glide.

Both now feed one glide on the listener tween key, which is the only
writer while the listener moves. The second message for a step
retargets that glide instead of starting a competing one, and it runs
at the speed the step itself carries. An EntityMove for the listener
that arrives alone still moves the listener, now per frame.
A LiveKit speaker with no position was placed at [0, 0, 0]: the centre
of the listener's current room at floor height. A phone-call partner, or
a speaker still connected from the room just left, was heard as a point
there, getting louder and swinging around as the listener walked.

A voice is now positional only while its speaker is an entity in the
current scene. The bridge passes null for anyone else, and MediaVoices
plays that voice as a stereo source at centre pan: it has no 3D panner,
so no position, no distance model and no direction. Cacophony fixes a
stream sound's panning mode at creation, so when the entity appears or
leaves, the voice's sound is rebuilt in the other mode on the same
track, keeping the named route it had or still wanted. A voice that
becomes positional is placed before it plays.

The test that asserted origin placement is rewritten to the new rule.
A Play for a playing key recreated its voice whenever start differed,
because the segment comparison included start. start is the join
cursor, not part of the region. The server re-Plays every audible sound
at its current cursor to a listener who enters a room, so every
continuing sound (music, a radio, a fire) was cut and restarted at
every doorway.

Region identity is now the repeat window, loopStart..finish, as the wire
contract says: same key, source and region keeps the voice, and an
explicit start seeks it. The MOO sends loopStart only for a sound that
repeats, so a single pass without it is kept when the voice's region
reaches back to the new cursor; a repeating sound without loopStart
still has a repeat window that follows start.

On a kept voice the seek is skipped when the playhead is already within
MEDIA_SEEK_TOLERANCE_MS (150 ms) of the cursor, measured across the loop
point for a looping voice, so a re-Play at the cursor the voice is at is
seamless. An engine that reports no playhead seeks as before, and
continue: false still forces a new voice.
A source's panning mode is fixed when Cacophony creates it, and nothing
handled a later change of is3d. The server sends is3d: false with no
position for a sound heard through a doorway that has no known point,
and is3d: true with a position once the listener is in its room. A kept
voice stayed in the wrong mode: still HRTF at its last coordinates, or
stereo and attenuated by distance with no direction. Keeping voices
across re-Plays makes this the common path at every doorway.

A Play or an Update that changes is3d for a playing key now builds the
same source and region in the new mode. The old voice plays on while
the replacement loads; at the swap the new voice starts at the old
one's playhead, with no second fade-in, unless the Play names a cursor
further than the seek tolerance away. Volume, gain, pitch, spatial
profile, loop count, occlusion and routing carry over: an Update's
fields are now folded into the payload that describes the sound, so a
rebuild (this one, or a moved segment) replays the state in effect now
rather than the original Play's. A Stop or a newer Play during the load
wins, and the replaced voice is released on every exit.

Ambisonic sounds are not rebuilt here, and a preloaded sound that was
never Played is left as it was.
@ctoth
ctoth merged commit 0ae1588 into feat/media-occlusion Oct 5, 2026
2 checks passed
@ctoth
ctoth deleted the fix/spatial-rendering branch October 5, 2026 03:36
ctoth added a commit that referenced this pull request Oct 5, 2026
Land the spatial rendering fixes (PR #164) on master
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant