1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
|
# asist-p — local voice assistant
[Español](README.es.md)
It listens through the microphone, thinks and answers out loud. Everything
runs on the machine; only the optional web search goes online.
It joins three existing engines (**Canary** to hear, **llama.cpp** to think
and **qwentts.cpp** to speak) in a Rust pipeline of threads and channels
where no stage waits for the next.
```
microphone ──samples──> segmenter ──utterance──> ASR ──text──┐
(cpal, real time) (VAD, turns) (Canary) │
v
speaker <──samples── synthesis <──sentences── conversation <─┘
(cpal, ring) (qwentts) (llama.cpp + tools)
```
Each box is a thread; each arrow, a channel. Two decisions carry the rest:
- **Nothing blocks the stage ahead.** The microphone never waits for the
recognizer, and the model never waits for the synthesizer. Whoever has
spare capacity drops work instead of accumulating delay.
- **Answers go out per sentence, not per answer.** Each sentence goes to the
synthesizer as soon as it closes, so the assistant starts talking while the
model is still writing. That is what separates half a second from five.
The assistant speaks Spanish: its prompts, tool descriptions and spoken
acknowledgements are in Spanish on purpose. The code is in English.
## Getting started
```bash
git clone --recursive https://git.all.ar/pub/asist-p.git && cd asist-p
scripts/bootstrap.sh # submodules, patches, binaries and models
cargo run --release -- check # check that everything is in place
cargo run --release -- run # start talking
```
`bootstrap.sh` is idempotent and does not copy models: it links them from
wherever they already are (`$ASIST_MIRROR`, default `~/GIT-MIRRO`, and
`$ASIST_HF`, default `~/HF`). If it cannot find the engines built, it builds
them, and that does take a while.
```bash
cargo run --release -- devices # list audio devices
cargo run --release -- run --barge-in # allow interrupting the assistant
cargo run --release -- run --shell # enable system commands
cargo run --release -- run --no-manage # do not start the servers
```
While developing, keep the servers running separately (loading the models
takes more than a minute) and restart only the Rust binary:
```bash
scripts/servers.sh start
cargo run --release -- run --no-manage
```
## Layout
| Crate | What it does |
|---|---|
| `asist-core` | Configuration, events, errors, HTTP client, telemetry and the tool registry |
| `asist-audio` | Capture and playback with cpal, voice detection, splitting into turns |
| `asist-asr` | Canary on ONNX: sliding window and final transcription |
| `asist-llm` | llama-server: SSE streaming, history, tool calls |
| `asist-tts` | qwentts: streaming synthesis and cloned voices |
| `asist-tools` | Web search, looking through the camera and at the screen |
| `asist-app` | Process supervisor, orchestrator and the `asistente` binary |
The three engines live in `vendor/` as submodules pinned to a specific
commit. The local changes on them are in `vendor/patches/`, and
`bootstrap.sh` applies them: without them the recognizer does not behave like
the one that was measured.
## Configuration
Everything is in [`config/asistente.toml`](config/asistente.toml), with
comments. Values that carry a number in their comment come from a specific
measurement, not an assumption; the details are in
[docs/RENDIMIENTO.md](docs/RENDIMIENTO.md) (in Spanish).
The most tweaked settings:
```toml
[vad]
silence_hold = 0.8 # how much silence ends your utterance
barge_in = false # interrupt the assistant by talking over it
[general]
history_turns = 8 # conversation memory
system_prompt = "..." # personality and style
[tools]
shell = false # run system commands
```
## Interrupting and speaker echo
By default the assistant is **half duplex**: while it talks, the microphone
is closed. That is not laziness: with open speakers the microphone hears
itself, and the assistant transcribes itself and answers itself.
With headphones, `--barge-in` lets you talk over it to interrupt. The
threshold is raised (`vad.barge_in_factor`) so the echo is not enough but a
real voice is.
## Latency
A real turn, measured end to end:
| Stage | Time |
|---|---|
| Recognition (2.5 s of speech) | 200 ms |
| Model, first token | 1660 ms |
| Model, rest | 1362 ms |
| Tools | 1 ms |
| **From the end of your sentence to the first audio** | **3213 ms** |
The binary measures this on every turn and flags the slowest stage. A few
configuration settings were worth more than any code change:
- The TTS decoded in 24 s blocks: **4948 ms → 585 ms** to first audio.
- The model template left `<think>` open: **8630 ms → 413 ms** to first token.
- The style prompt suppressed tool calls: **0/8 → 8/8**.
- The camera captured at 720p: **7.8 s → 2.9 s** per image at 640x480.
- Tools ran in silence: **~8.4 s → 4.1 s** until something is heard.
- The screen at 640 px made the text up: **1 of 3 right → 3 of 3** at 1280 px.
All of them, with the reasons and how they were measured, are in
[docs/RENDIMIENTO.md](docs/RENDIMIENTO.md) (in Spanish).
## Tools
The assistant can call functions. It ships with these:
| Tool | What it does | Measured cost |
|---|---|---|
| `hora_actual` | System date and time | 0 ms |
| `buscar_en_internet` | Searches and summarizes | 0.5–3.6 s |
| `mirar_por_la_camara` | Takes a photo and answers about what it sees | 4.1–5.2 s |
| `mirar_la_pantalla` | Captures the screen and answers about what is on it | 7.6–12.4 s |
| `ejecutar_comando` | System commands, **off** by default | — |
The tool names are Spanish because the model reads them in a Spanish
conversation.
While a tool works, the assistant says «Déjame que lo busque» or «Voy a
mirar». It is not decoration: without it the turn spends six seconds in
silence and reads as a hang. With the acknowledgement, the first audio
arrives in 3–4 s.
**Search** supports three backends: `tavily` (with a key, returns an already
written answer), `ddgs` (no key, the same library behind `duckduckgo-mcp`)
and `searxng`. The command backend runs a script that writes JSON, so adding
another engine means writing a script, not touching Rust.
**Vision** (camera and screen) takes advantage of the server already loading
the multimodal projector: the same model that converses describes the image.
Neither the photo nor the capture is saved to disk, and they do not enter the
history.
The screen uses 1280 px and the camera 640, and the difference matters: a
scene can be understood, but text has to be read. Measured, at 640 px the
model does not say it cannot read the screen, **it makes up what it says**
(it answered that the meeting was «at 10:00» when it said 15:30). Hence the
threefold cost.
`ejecutar_comando` is off on purpose: giving a shell to a model that obeys
what it hears through the microphone is a change of security posture. When it
is on, only allowlisted commands get through, with no shell to interpret
metacharacters and with a deadline.
Adding your own takes about twenty lines: [docs/EXTENDER.md](docs/EXTENDER.md)
(in Spanish).
## Voice
The repository ships no cloned voice: a voice belongs to someone. To make the
assistant speak with one, record a few seconds of audio with its transcript
and extract the latents:
```bash
scripts/clone-voice.sh recording.wav transcript.txt
```
They end up in `assets/voices/` (outside git). Without a cloned voice, remove
the `[tts.reference]` section from `config/asistente.toml` and the assistant
uses one of the model's own voices.
## Tests
More than a hundred unit tests that need neither models nor a microphone,
plus integration tests against the real servers.
```bash
cargo test # unit tests, no models
scripts/servers.sh start
cargo test --release -p asist-app --test integration -- --nocapture
cargo test --release -p asist-app --test integration -- --ignored # full loop
```
The integration tests skip themselves when no server is listening, and they
take turns on the GPU through a mutex: both servers share a 4 GB card, and in
parallel they measure contention instead of latency.
The `--ignored` test closes the loop without a microphone: it synthesizes a
sentence and checks that the recognizer gets it back.
## Requirements
- Rust 1.85 or later
- CUDA for the C++ engines (they run on CPU, but barely)
- PipeWire or ALSA
- ~6 GB of disk for the models
- `ffmpeg` for the camera and to scale down captures
- For the screen: `grim` (Wayland) or `maim`/ImageMagick (X11)
- For search: a Tavily key in `$TAVILY_API_KEY`, or `uv tool install ddgs`
## License
MIT
|