Voice agents that handle a call on their own and stop before anything that matters. Small enough to run on a CPU, so the audio never has to leave the building.
Most voice products ask you to choose: let it act, or listen to every call. The interesting line is narrower than that. Let it act on the ninety percent that is answering questions, and make it stop on the ten percent that spends money or changes a record.
An agent takes a call on its own. Watch what happens when it reaches the part that spends money.
A voice model does not need a datacentre to answer a phone. Ours are built to fit on ordinary CPUs, which changes three things at once — and the third is the one that closes deals.
The unit economics of a phone line stop being an inference invoice. That is what makes it sellable to a dental practice rather than only to an enterprise.
Latency you can hear is latency that loses the caller. Running next to the phone system removes the round trip that causes it.
Nothing leaves the building unless you send it. For anyone handling health, legal or financial calls, that is not a feature — it is the reason they can say yes at all.
Small models get better by being taught, not by being replaced. A reading agent works through new papers and open source on its own, keeps the few findings that would change how the voice models are trained or served, and leaves the rest. The same discipline as the phone: act freely on the gathering, stop before the change.
Everything new in speech, quantisation and small-model training.
The overwhelming majority. A finding that changes nothing is noise.
A specific change, with the reason and the expected cost.
A person approves it before it touches a model in production.
We are taking a small number of first lines — the kind where a wrong answer actually costs something, because that is where the guardrail has to earn its place.
One reply from a person, not a sequence.