Showing live voice agent transcripts in a Flutter app with LiveKit
If you build a voice agent with LiveKit Agents and a Flutter frontend, at some point you want the conversation on screen: what the user said and what the agent is saying, updating live. The docs cover it, but the pieces are spread out and the older API in many examples is deprecated. Here's what works with the current SDKs. Where the text comes from When AgentSession runs STT, it publishes transcriptions to the room on the lk.transcription text stream topic. The agent's own speech goes out on the same topic, synced with audio playback, so it can show up word by word as the agent talks. If the user interrupts, the agent text is cut to what was actually spoken. The sender identity on each stream is the participant who was transcribed, so you can tell user and agent apart without extra metadata. The older TranscriptionReceived event and publish_transcription() are deprecated. They use a separate delivery path, so anything written to lk.transcription never reaches a TranscriptionEvent listener. Build on text streams. The Flutter handler import 'dart:convert'; final transcript = {}; void listenForTranscripts(Room room) { final me = room.localParticipant?.identity; room.registerTextStreamHandler('lk.transcription', (TextStreamReader reader, String participantIdentity) async { final info = reader.info!; if (info.attributes['lk.transcribed_track_id'] == null) return; // chat message, not a transcript final segmentId = info.attributes['lk.segment_id'] ?? info.id; final speaker = participantIdentity == me ? 'You' : 'Agent'; var text = ''; void update(bool isFinal) { transcript[segmentId] = (speaker: speaker, text: text, isFinal: isFinal); onTranscriptChanged(); // setState, notifyListeners, etc. } try { await for (final chunk in reader) { text += utf8.decode(chunk.content); update(false); } } catch (_) { return; // stream aborted, e.g. during a reconnect } // read this after the stream closes: agent streams set it in the trailer update(info.attributes['lk.transcription_final'] == 'true'); }); } The docs example uses reader.readAll() . That waits for the stream to close, and an agent segment stays open until the agent stops talking, so the whole reply lands at once. Reading chunks gives you the word-by-word version. Interim vs final User speech and agent speech arrive differently. For user speech you get a new stream for each interim result, then a final one, all with the same lk.segment_id . Key your map by that id and replace the entry every time, or the same sentence shows up three times. Agent speech is one stream per segment that grows as the agent talks. It opens with lk.transcription_final set to false and flips it to true in the trailer when it closes, which is why the handler checks it after the loop. If you only want finished sentences, skip entries where isFinal is false. Register once registerTextStreamHandler throws if a handler for that topic is already set. In a widget that can rebuild or reconnect, register once after connecting and call room.unregisterTextStreamHandler('lk.transcription') in dispose . If the transcript still feels laggy after this, look at end-of-turn detection on the agent side. I wrote about tuning that in turn detection and barge-in for voice agents, and if you are estimating what a production voice agent costs per minute, there's a voice AI cost calculator on our site. The code was checked against the client-sdk-flutter source and the LiveKit docs. Originally published on Axionry Engineering. Top comments (0)
Comments
No comments yet. Start the discussion.