fix(voice): push voice presence to user on mid-session space join

Voice presence (voiceStates/voiceUserStates/spaceVoiceStates) was only ever
delivered in the WS `ready` payload — i.e. at connect/reload. A user joining a
space mid-session got `member_joined` (no voice data) and a bare space object;
`GET /api/spaces/:id` (the channel-sidebar hydrator) carries no voice state
either. So members already sitting in a voice channel stayed invisible in the
new member's sidebar until a full page reload.

Fix at the systemic root: ConnectionManager.addUserSpace — the single chokepoint
every join path funnels through (invite, public join, join-request approval),
and which is NOT used on reconnect (that path uses setUserSpaces) — now pushes a
scoped `space_voice_state` snapshot to the joining user. The snapshot is built by
a new buildSpaceVoiceState(spaceId, userId) helper that is also the single source
of truth feeding buildReadyPayload (refactored to use it), so the connect-time
and join-time paths can never drift.

Robustness:
- Delivered over the same ordered WebSocket as voice_state_update deltas — no
  REST snapshot-vs-event-stream race.
- VIEW_CHANNEL-filtered via computePermissions exactly like `ready`: a joiner is
  never told who occupies a voice channel they cannot see.
- Client applies it scoped to the space (utils/voiceStateSync.applySpaceVoiceState):
  merges occupants/statuses and rebuilds only that space's restriction keys,
  never disturbing voice state in other spaces.
- Skipped when the space has no active voice and no restrictions (e.g. space
  creation).

Tests: server helper behavior, the join push, and private-channel exclusion;
client scoped-apply. Specs updated (websocket.md, voice.md, spaces.md).
This commit is contained in:
Jannis Braun
2026-06-30 17:00:37 +02:00
parent f807524103
commit e84daf57aa
9 changed files with 510 additions and 42 deletions
+12
View File
@@ -17,6 +17,18 @@ Source files:
5. Client calls `POST /api/livekit/token { channelId }` → gets JWT + LiveKit URL
6. Client connects to LiveKit room with token
### Voice presence bootstrap on mid-session space join
A client learns who is sitting in a space's voice channels from the WS `ready`
payload at connect time. Joining a space *without reloading* therefore needs the
same bootstrap for the new space, or its voice channels render empty until a
refresh. The server pushes a scoped `space_voice_state` snapshot from
`ConnectionManager.addUserSpace` (the single join chokepoint), built by
`buildSpaceVoiceState(spaceId, userId)` — the same VIEW_CHANNEL-filtered helper
that feeds `ready`. The client applies it via `utils/voiceStateSync.applySpaceVoiceState`.
See `docs/systems/websocket.md` → "Mid-session space join" for the full rationale
(single ordered channel, no snapshot-vs-stream race).
### Microphone pre-arm (iOS user-gesture discipline)
`utils/voice.joinVoiceChannel` fires `AudioContext.resume()` and `AudioManager.setInputDevice(inputDeviceId)` (which ends in `getUserMedia({audio:…})`) **synchronously inside** the click handler, before the `connectFn(channelId)` call. iOS Safari only surfaces the microphone permission prompt when `getUserMedia` is invoked from inside an active user-gesture; the original flow only acquired the mic in `useLiveKit`'s `syncMic` effect, which fires AFTER `room.connect()` resolves (token fetch + WS handshake) — many awaits past the gesture window. iOS PWA standalone is especially strict and would silently never surface the prompt; the user would see "Waiting for others to join…" indefinitely until they locked/unlocked the device (which iOS treats as a fresh activation).