Google Gemini Live Agent Challenge Technical Entry
Live streaming is the new frontier of global commerce, yet solo creators often lack the support of a professional production crew. StreamCoach Live was engineered to bridge this gap, serving as a real-time "Digital Director" that sees, hears, and guides streamers. To make this possible, we had to move beyond simple API calls and design a high-performance, low-latency architecture capable of full-duplex multimodal reasoning.
Technical Table of Contents
- 1. Inspiration: Elevating the Solo Creator Experience
- 2. Architecture Overview: The Four Pillars of StreamCoach
- 3. The Engineering Stack: Why Go and AudioWorklet?
- 4. Solving Real-Time Conflicts (Feedback & Deduplication)
- 5. Global Impact: Democratizing High-End Sales Coaching
- 6. Conclusion: The Road to OBS and Beyond
1. Inspiration: Elevating the Solo Creator Experience
Live streaming is a high-stakes performance where every second of silence counts. Streamers on platforms like TikTok and YouTube must juggle product demonstrations, audience chat, and technical monitoring simultaneously. Post-stream analytics are useful, but they don't help a streamer who is losing viewers right now.
We envisioned a companion that could provide "tactical whispers"—instant, proactive cues whispered in the streamer's ear. By leveraging Vertex AI, we transformed the AI from a simple responder into an active participant in the stream's success.
2. Architecture Overview: The Four Pillars of StreamCoach
StreamCoach Live utilizes a sophisticated full-duplex architecture designed specifically for ultra-low latency and complex multimodal reasoning. We divided the system into four distinct phases to handle the constant flow of video and audio data.

Phase 1
High-Fidelity Multimodal Capture
The system begins at the edge with a custom-built Chrome Extension (Manifest V3). Unlike standard screen recording tools, our extension acts as a specialized multimodal sensor that isolates and optimizes data before it even reaches the cloud.
- Visual Grounding: We capture 1280x720 (HD) frames at a steady 1 FPS. This resolution was carefully chosen as the "sweet spot" for AI Vision; it is sharp enough for Gemini to read tiny scrolling chat text and identify fine product details, yet efficient enough to prevent browser lag.
- The AudioWorklet Breakthrough: Standard browser audio recording is often prone to stuttering under heavy CPU load (common when running TikTok tabs). We implemented the Web Audio API’s AudioWorklet to process raw 24kHz PCM audio in a dedicated background thread. This ensures the streamer's voice is captured with professional clarity, completely isolated from the stream's music or host audio.
Phase 2
Google Cloud Orchestration
The "Heart" of StreamCoach is a high-performance Golang Orchestrator deployed on Google Cloud Run. It serves as the bridge between the browser and the AI brain, managing high-frequency bidirectional traffic.
- WebSocket Multiplexing: The orchestrator maintains a persistent, full-duplex link. It multiplexes visual frames, audio chunks, and system instructions into a unified stream for the AI.
- Thread-Safe Concurrency: To handle the rapid-fire nature of multimodal data, we implemented strict
sync.Mutexlocks. This prevents "Concurrent Write" panics, ensuring the media stream remains stable and uninterrupted even during intense interactions. - Dynamic Cloud Authentication: We leveraged Google's Application Default Credentials (ADC). Our backend intelligently switches between local development tokens and production-grade Service Account credentials on Cloud Run, ensuring secure and seamless access to Vertex AI resources.
Phase 3
Gemini Live 2.5 Flash Native Audio Intelligence
The "Brain" of the operation is Vertex AI’s Multimodal Live API (LlmBidiService). This is where the magic happens—moving beyond text prompts into the realm of real-time intuition.
- Simultaneous Reasoning: Gemini 2.5 Flash is uniquely capable of processing interleaved video and audio chunks as they arrive. This allows the AI to provide feedback based on the synergy of what it sees and hears, rather than treating them as separate data points.
- Visual Directives: Because the AI has HD visual access, it can specifically answer questions like "What is the host holding?" or "Is the lighting too dark?". It provides specific, grounded answers rather than generic advice.
- Native Audio Generation: We bypassed traditional Text-to-Speech (TTS) bottlenecks. Gemini generates native audio speech, conveying natural human intonation and supporting local Indonesian dialects ("gaul"). This makes the "Coach" feel like a supportive human partner rather than a robotic assistant.
Phase 4
The Interleaved Feedback Loop
The final phase is the delivery of actionable intelligence back to the streamer in a way that feels non-intrusive yet impossible to miss.
- Bimodal Delivery: Feedback is interleaved. Streamers receive "Whispers" (AI Voice Feedback) in their ear and "Coaching Cards" (AI Insight Recommender) on a dynamic dashboard simultaneously.
- Smart Audio Ducking: To ensure the user hears every cue, the extension implements an intelligent ducking system. When the AI starts to speak, the live stream volume is automatically lowered to 20%, returning to 100% only after the AI has finished its tactical suggestion.
- Push-to-Talk (PTT) Synchronization: We built a robust PTT system that allows the streamer to take control. Pressing the button immediately interrupts the AI and mutes the stream, creating a focused channel for the user to ask the "Boss" questions.
3. The Engineering Stack
Our stack was selected for its ability to handle "Streaming at the Edge" while scaling effortlessly in the cloud:
- Backend: Go (Golang) - chosen for its superior handling of concurrent WebSockets and memory efficiency.
- Infrastructure: Google Cloud Run & Artifact Registry - providing a serverless, cost-effective, and scalable environment.
- AI Core: Vertex AI Gemini 2.5 Flash Live - the most advanced multimodal API for low-latency reasoning.
- Frontend: Manifest V3 Chrome Extension - utilizing AudioWorklet for dedicated audio processing threads.
4. Hard-Won Engineering Lessons
Building a real-time agent presented unique challenges. One major hurdle was the Feedback Loop: the AI would hear its own voice through the microphone and respond to itself. We solved this by implementing Exclusive Mic Mixing, where the system audio is muted at the hardware level whenever the user speaks. Additionally, we developed an Anti-Spam Deduplication Filter in Go to prevent the AI from repeating identical responses, ensuring a clean and professional interaction every time.
5. Social Impact: Democratizing Production
StreamCoach Live has a mission beyond technology. We aim to Democratize Professional Sales Coaching. In the past, only massive enterprises could afford a professional director. By giving MSMEs (UMKM) access to high-end coaching at near-zero cost, we empower small businesses to compete with global brands. Furthermore, it serves as a tutor for Soft-Skill Development, helping creators master public speaking, eye contact, and articulation—skills vital for the modern workforce.
6. Conclusion: The Road Ahead
We are currently working on expanding StreamCoach Live into a dedicated OBS Plugin to support professional gaming setups. Our future roadmap also includes real-time Sentiment Analysis to track viewer emotions and Automated Highlight Clipping, making StreamCoach the ultimate end-to-end tool for the modern digital creator.
Disclosure: This project and technical deep-dive were created for the purpose of entering the Gemini Live Agent Challenge hackathon. We are proud to build on Google Cloud and showcase the power of Gemini 2.5 Flash. #GeminiLiveAgentChallenge

Comments