move from private r&d repo
This commit is contained in:
@@ -0,0 +1,34 @@
|
||||
# Conversation
|
||||
This proof of concept is for implementing conversational loop with AI at a low level with raw C.
|
||||
|
||||
## How it works
|
||||
This simple program:
|
||||
1. Captures audio samples for x seconds into a local file using alsa
|
||||
2. Sends the raw pcm data to Deepgram STT API for transcription
|
||||
3. Uses the transcript to call Anthropic to get an assistant's response
|
||||
4. Converts Anthropic's response to audio using Deepgram's TTS
|
||||
5. Plays back the generated audio from Deepgram using alsa
|
||||
6. Quits. _There is no conversational loop at the time of this writing._
|
||||
|
||||
## Requirements
|
||||
Platform: the project is meant for Linux x64. It was successfully run on arch linux, with no guarantees otherwise.
|
||||
|
||||
* libcurl
|
||||
* build-essential
|
||||
* libasound2-dev
|
||||
|
||||
_See more deps in Makefile `install-deps` command._
|
||||
|
||||
## Getting started
|
||||
1. Add `DEEPGRAM_API_KEY` and `ANTHROPIC_API_KEY` to `.bashrc` as environment variables.
|
||||
2. Build & run
|
||||
```sh
|
||||
make
|
||||
./main
|
||||
```
|
||||
|
||||
## Audio configuration
|
||||
Please see the `audio.h` and `audio.c` as the params for alsa were hardcoded based on my local arch linux x64 machine. You may have to adjust it based on your setup (e.g. you may have a mic with 2 channels/stereo, which would change the size of each frame).
|
||||
|
||||
## Future work
|
||||
Add an event loop, maybe add multi-threading, so that we can interrupt the AI, ask follow-up questions, and make it sound like a back and forth conversation.
|
||||
Reference in New Issue
Block a user