VoxNote: Real-Time Call Transcription with AI Insights
Real-time speech-to-text transcription system with speaker diarization and GPT-powered insights, built on AWS with Kafka streaming.
The Challenge
Businesses needed a way to convert live conversations into actionable intelligence. Existing transcription tools lacked real-time speaker identification, and couldn't generate meaningful summaries or follow-up suggestions during the call itself.
Our Approach
We engineered a full-stack real-time transcription platform combining state-of-the-art speech models with streaming infrastructure and generative AI.
Real-Time Transcription Pipeline
OpenAI Whisper handles audio-to-text conversion with high accuracy across accents and domains. The DIART algorithm performs real-time speaker diarization — identifying who is speaking at any moment — enabling accurate multi-speaker transcripts.
Streaming Architecture
Amazon MSK (Managed Streaming for Apache Kafka) handles the high-throughput data flow between microservices. Audio chunks are streamed from the client, processed through transcription and diarization, and delivered back with minimal latency.
AI-Powered Insights
GPT integration analyzes transcripts in real-time, generating conversation summaries, action items, sentiment analysis, and contextual suggestions that surface during the call.
Frontend Experience
A React.js frontend leverages the Web Audio API for browser-based audio capture and streaming, with a clean interface showing live transcripts, speaker labels, and AI insights.
Technical Stack
- Cloud: AWS (EC2 with GPU, MSK, S3)
- Speech: OpenAI Whisper, DIART diarization
- AI: GPT for summaries and insights
- Streaming: Apache Kafka via Amazon MSK
- Frontend: React.js, Web Audio API
- Backend: Python
Results
- Real-time transcription with accurate speaker identification
- AI-driven insights delivered during live conversations
- Scalable streaming architecture handling concurrent sessions
- Actionable summaries enabling smarter business decisions

