Conversation
This proof of concept is for implementing conversational loop with AI at a low level.
- Capture audio
- Use deepgram for live transcription
- Detect final speech/turn detection
- AI answer -> text
- Text -> generated voice with deepgram
- Playback on client side
Requirements
Platform: the project is meant for Linux, specifically Ubuntu desktop.
Implementation
The loop should be as follows:
- create the thread that manages the deepgram websocket client
- grab audio and send over to other thread to process & send
- stream pcm packets over the websocket connection
- print inbound messages from websocket connection to stdout
- signal in stdout when there is a "final message", which signifies "turn detection"
Open questions
- Can deepgram websocket stream take any format? raw pcm? Answer here
- Consider using Golang to interop with c, whereby the golang uses deepgram sdk, and the c library for alsa.