CHITRAGUPTA v2 · HOW IT WORKS
One document, two models, and arithmetic in between
The world document is primary state; speech is a side-effect. Everything on this page is a consequence of taking that literally — including the parts that made the system harder to build.
What actually happens over a session
Read the two middle columns against each other. The document is written on almost every row; speech happens on very few. That gap is the architecture — and the version before this one had no such gap, because deciding what to say was the job of every frame.
There is no open microphone. Voice input is push-to-talk. A turn only ever begins from one of three things — and two of them run without anyone touching the phone, which is the entire point.
The poll exists for two reasons at once. It is what lets a find announce itself even if the camera has gone dark since — the trigger tests the document, not this tick's caption — and on a free host it is also what keeps the server from falling asleep mid-task.
Where the lock is, and where it is not
The document is one JSON file, so load → mutate → save must be atomic and there is exactly one lock. Originally it was held across the whole tick, which meant a tick owned the document for two model round trips plus any web search they made. Measured at the point where a question arrives mid-reasoning, that cost the user two full seconds of nothing.
Two consequences follow, and both are load-bearing. Every window reloads from disk on entry — a window that reuses a document read before a model call silently rolls back whoever wrote in the meantime, a lost update that leaves no trace because both writes reported success. And the document's revision number is bumped on window entry, not on save: windows are serialized by the lock, so entry order is write order, and the browser can drop any render older than one it has already painted.
A server-side concurrency fix is not done until the client can exercise it.
None of the above reached a user for some time, because the browser held a single busy flag covering both ticks and chat. A typed question was queued client-side with a status line apologising for a constraint that no longer existed anywhere in the system. The message never left the browser. Two flags now protect genuinely different things: frames must not stack, and replies must not interleave.
The document itself
There is no retrieval step and no memory tool. Every prompt is built by rendering the document, because a model that has to ask for its state will reliably forget to ask. The section order below is not cosmetic.
The order is stability-first. Title, tasks, narrative and environment facts change rarely; raw captions change on every tick. So the things that move least go at the top, and consecutive ticks share the longest possible unchanged prefix — which is what the reasoning provider's prefix cache is able to charge less for. Ordering the document by what a reader would find natural would have cost real money on every frame.
The caption buffer is bounded at 24. On overflow the oldest sixteen are summarised into the narrative by one cheap model call, with time spans preserved rather than only facts, plus any durable environment facts worth promoting. Raw captions are never silently dropped, and the freshest window is never compacted.
The trigger engine costs nothing, and it is the only thing that starts a sentence
Wall-clock arithmetic is free; tokens are not. The trigger check runs on every tick and every 20-second poll, costs nothing at all, and the reasoning model is only woken when it returns something — or when a frame arrived, or the user spoke.
| Trigger | Fires when |
|---|---|
| expectation_due | a time-anchored expectation passed its deadline unresolved |
| stale_task | an in-progress task has gone unmentioned for eight minutes |
| proposal_pending | a plan was proposed and unanswered for 150 seconds |
| wanted_found | a find-list item was seen and the user has not been told — coalesced into one event for all of them |
| wanted_stuck | a search hit its miss or unclear budget. Ask the user rather than failing silently |
Each one claims its own state on firing, before anything downstream awaits. An unclaimed trigger fires again on the next tick while the first one is still thinking — the same double-fire lesson the timer system taught in v1.
Unprompted speech is then gated by a 90-second politeness budget, and exactly two things bypass it. [URGENT] — physical risk, or work about to be ruined — and the follow-up window that a user's own question opens.
Speaking used to reset the politeness budget. Asked to find the onions, the assistant replied “I'll point them out as soon as they're in view” — and that reply gagged it for the entire 90-second search. It found them at +27 seconds, logged them silently, and said nothing until it was asked again.
Speaking should make the assistant quieter. Being asked should make it more forthcoming. Those pull in opposite directions, so they are tracked as two separate timestamps.
The find list, and the door that was closed
“Find the chicken and the onions” opens a search. Every open item is put to the camera by name, in one block, on every frame — one block rather than one question per item, so the list never competes with the four-watch cap and a fourth search cannot silently stop reaching the camera.
chicken packet: FOUND — middle shelf, behind the milk onions: NOT VISIBLE — this part of the frame is the door rack
There is no tool that marks an item found. The model can open a search and cancel one; only plain string matching over the caption — zero tokens — sets the found state. That is what stops an inference standing in for an observation. The caption said “several bags of lentils” and the model once upgraded it to “I can see the beans”. The found state now has no model-facing door.
Two smaller decisions here both came from thinking about what happens when something goes wrong. Found and announced are separate flags— one is about the world, the other about speech — so a failed utterance simply re-fires next tick instead of leaving an item found and unspoken forever. And matching is by name, never by question index: the index is positional over a list rebuilt each tick while the vision call runs with the lock released, so eventually it would announce one item's location under another item's name.
What one minute of watching costs
In this split the vision call is the only image cost, so its prompt tokens are the per-frame bill — which is what makes the number knowable at all. A tick is roughly 1,350 input and 90 output tokens at the coarse tier; the close-up tier caps the longest side at 1,024 px, which is about a thousand image tokens on its own.
| Free tier, 8k/min · 200k/day | DeepInfra, metered | |
|---|---|---|
| fastest sustainable tick | one every ~11 s | no per-minute ceiling |
| a whole day's allowance | ~139 ticks, total | about four cents |
| 1,000 ticks | not reachable | ~$0.26 |
v1 survives on a free tier because its cadence is slower and most of its turns are text-only. v2 ticks continuously by design — it is a fundamentally heavier vision consumer, and the honest conclusion was that the free tier is not a constraint it can be engineered inside. That is now enforced rather than documented: the wrong provider refuses to start.
The close-up tier gets a nudge, then a backstop. The count of close-up frames is rendered into the document, so the drift is visible to the thing causing it, and a hard cap forces coarse at 120. v1 proved that a fine mode set once is never voluntarily reverted; v2's own live session proved the opposite failure, and the tier was never chosen at all.
Fifteen tools, and two that are deliberately missing
| Tool | What it does |
|---|---|
| propose_plan · commit_plan · discard_plan | The approval cycle. A proposal is spoken, never written to tasks, and nothing tracks it until the user agrees |
| add_wanted · drop_wanted | Opens and cancels a search. Neither one can mark an item found — only the camera can do that |
| update_tasks · mark_task | For a plan the user is already working through, rather than one being proposed |
| set_expectation · resolve_expectation | A deadline, or a condition put to the camera verbatim on every frame until it resolves |
| set_vision_focus | The standing lens. Replaced, never appended, so duplicates are impossible by construction |
| log_environment · retract_environment_fact | Durable spatial memory, and undoing it. The retraction takes a correction, not just a deletion |
| web_search · fetch_page · calculate | Inherited from v1 unchanged, and flagged blocking so they never run on the event loop |
The document-mutating tools close over the agent's current in-memory document, so a turn's tool calls and the agent's own writes can never interleave on disk. And retract_environment_fact takes a correction, not just a deletion — the raw captions that produced the wrong inference are still in the buffer and will suggest it again on the very next tick. A hole in the fact list does not block that; a durable “the bag on the pantry shelf is NOT toor dal” does.
Four designs I rejected
| Considered | Rejected because |
|---|---|
| A VideoLLM-online streaming architecture | Examined seriously as a route to real-time. Its pooled per-frame vector is an unconditional compression that cannot preserve a label or a torque figure, and its cache grows without bound so latency degrades over a session. The lesson taken instead: speech latency is not a real ceiling, because synthesis can start before the next frame is processed |
| A duplicate-watch check | Built, measured, removed. The duplicates that actually occurred shared about 19% of their words — “car is safe to work under” versus “the car should be settled firmly on both ramps” — while a threshold low enough to catch those merged genuinely distinct watches differing only in the object. Silently dropping a watch the user is relying on is far worse than carrying a duplicate |
| A hard cutoff on stale searches | A search past its budget now stays open but stops buying full resolution, and the prompt gets a nudge to close it. Dropping it outright would silently abandon something the user may still be waiting on |
| A multi-agent orchestrator | Rejected for v1 and the reasoning holds. The reasoning model orchestrates itself — it reads the document, decides mid-thought whether to call a tool, and routes to the right response. The thinking chain is the orchestration |
The duplicate check is the one I would highlight. It was a reasonable idea, it was built, and measuring it is what killed it — in both directions at once. Removing it also removed the pressure it existed for, because form and safety moved into a tool that replaces rather than appends, which makes duplicates impossible by construction instead of by threshold.