35 lines
1.4 KiB
Markdown
35 lines
1.4 KiB
Markdown
# Conversation
|
|
This proof of concept is for implementing conversational loop with AI at a low level with raw C.
|
|
|
|
## How it works
|
|
This simple program:
|
|
1. Captures audio samples for x seconds into a local file using alsa
|
|
2. Sends the raw pcm data to Deepgram STT API for transcription
|
|
3. Uses the transcript to call Anthropic to get an assistant's response
|
|
4. Converts Anthropic's response to audio using Deepgram's TTS
|
|
5. Plays back the generated audio from Deepgram using alsa
|
|
6. Quits. _There is no conversational loop at the time of this writing._
|
|
|
|
## Requirements
|
|
Platform: the project is meant for Linux x64. It was successfully run on arch linux, with no guarantees otherwise.
|
|
|
|
* libcurl
|
|
* build-essential
|
|
* libasound2-dev
|
|
|
|
_See more deps in Makefile `install-deps` command._
|
|
|
|
## Getting started
|
|
1. Add `DEEPGRAM_API_KEY` and `ANTHROPIC_API_KEY` to `.bashrc` as environment variables.
|
|
2. Build & run
|
|
```sh
|
|
make
|
|
./main
|
|
```
|
|
|
|
## Audio configuration
|
|
Please see the `audio.h` and `audio.c` as the params for alsa were hardcoded based on my local arch linux x64 machine. You may have to adjust it based on your setup (e.g. you may have a mic with 2 channels/stereo, which would change the size of each frame).
|
|
|
|
## Future work
|
|
Add an event loop, maybe add multi-threading, so that we can interrupt the AI, ask follow-up questions, and make it sound like a back and forth conversation.
|