← Research

2026-09-22 · Engineering notes · Anamanti

Fast thoughts and slow thoughts

A voice assistant lives or dies on how quickly it answers, so Anamanti thinks in two speeds: a cheap reflex, and the slow, expensive brain behind it.


§ Fast thinking, slow thinking

Borrowing from Kahneman, Anamanti has a “System-1” reflex that answers easy requests (“set a timer,” “what’s the weather”) in one fast step, and only wakes the slow, expensive LLM when a question needs reasoning or memory. The reflex runs on a small local model (Laya-Decision), with a safety valve: a “does this actually need real understanding?” check that makes it hand off rather than guess.

§ Why that’s the right thing to speed up

The reflex is aimed at a specific, structural slowness. Because the assistant always has tools available, the full LLM path takes its slowest branch on every turn, and it also does a round-trip to look up memory before answering. A request the reflex can resolve skips both. It also, on purpose, leaves alone the parts that happen before you’ve finished talking, because there’s nothing there for it to help.

§ Talking sooner

For anything that does need the full brain, Anamanti speaks the first sentence while it’s still writing the rest, so the first audio arrives in about two seconds instead of eight. The catch was that the device ends its turn on the first “stop” signal, so all the little speech chunks are bundled into one continuous stream and it never cuts itself off mid-thought.

The safety valve is the part that makes it work. A classifier that can only pick from a list is dangerous in front of an open-ended assistant, unless it also knows how to say “this one’s above my pay grade” and defer.