Back to case studySource

CHITRAGUPTA v2 · HOW IT WORKS

One document, two models, and arithmetic in between

The world document is primary state; speech is a side-effect. Everything on this page is a consequence of taking that literally — including the parts that made the system harder to build.

01

What actually happens over a session

Read the two middle columns against each other. The document is written on almost every row; speech happens on very few. That gap is the architecture — and the version before this one had no such gap, because deciding what to say was the job of every frame.

USERTHE DOCUMENTSPEECHVISION0:00asks about dinnerpropose_plan · 6 stepsreads the plan aloudcaption0:20caption folded innothing saidcaption0:40—never asked✕ dropped · scene unchanged2:30proposal re-raised▸ trigger: proposal_pending“still waiting…”caption3:10“yes this style…”commit_plan · update_tasksset_vision_focuscommits, names step 1reused · unchangedlater — walking to the fridgecaption folded in“Slow down…”[URGENT] · gate bypassedcaption · motion blurlog_environment × 2nothing saidcaptionlater — the lens capcaption folded inreports the black view✕ flat frame · still billedcompaction · 13:24:56–13:36:58nothing saidcaption18:00—429 · daily cap hit✕ —filled dot = spoken aloud · open ring = the speech question was asked and answered “no” · blank = it was never asked
Eighteen minutes, and the user speaks three times. Note the three different reasons nothing was said: a tick that asked the speech question and answered no, a tick where the question was never asked because nothing had happened, and a frame the browser dropped before it became a request at all. Only one of those three costs anything. The last row is the session ending on a provider it should never have been able to reach.
WHAT IS NOT TRUE OF THIS DIAGRAM

There is no open microphone. Voice input is push-to-talk. A turn only ever begins from one of three things — and two of them run without anyone touching the phone, which is the entire point.

You ask
typed, or one press of the mic button
A camera tick
a frame, on an interval, gated in the browser
A poll
every 20 s — arithmetic only, and usually free

The poll exists for two reasons at once. It is what lets a find announce itself even if the camera has gone dark since — the trigger tests the document, not this tick's caption — and on a free host it is also what keeps the server from falling asleep mid-task.

02

Where the lock is, and where it is not

The document is one JSON file, so load → mutate → save must be atomic and there is exactly one lock. Originally it was held across the whole tick, which meant a tick owned the document for two model round trips plus any web search they made. Measured at the point where a question arrives mid-reasoning, that cost the user two full seconds of nothing.

ONE TICK1+2caption~1.5 s3afold in~1 ms3bStage 1 — bookkeeping~2 s3cworth asking~1 ms3dStage 2 — speech~1 s3egate~1 msLOCKreleasedreleasedreleasedEvery model call in the system happens in a released phase.The three locked phases are reload → mutate → save, and the document is one JSON file.Not to scale — a true scale would render the three locked phases as invisible slivers.
The rule the whole restructure exists to enforce: never await a model call inside a write window. A network call inside one re-serializes ticks and chat and silently undoes the split — the code still looks correct, it is just slow again. There is exactly one deliberate exception, compaction, because it rewrites two sections together and cannot be replayed against a document that moved underneath it.
a question reaching the wire — lock across the whole tick2.00 s
the same question — lock across writes only0.01 s

Two consequences follow, and both are load-bearing. Every window reloads from disk on entry — a window that reuses a document read before a model call silently rolls back whoever wrote in the meantime, a lost update that leaves no trace because both writes reported success. And the document's revision number is bumped on window entry, not on save: windows are serialized by the lock, so entry order is write order, and the browser can drop any render older than one it has already painted.

A server-side concurrency fix is not done until the client can exercise it.

None of the above reached a user for some time, because the browser held a single busy flag covering both ticks and chat. A typed question was queued client-side with a status line apologising for a constraint that no longer existed anywhere in the system. The message never left the browser. Two flags now protect genuinely different things: frames must not stack, and replies must not interleave.

03

The document itself

There is no retrieval step and no memory tool. Every prompt is built by rendering the document, because a model that has to ask for its state will reliably forget to ask. The section order below is not cosmetic.

[Current time]Stamped in the user's timezone, not the server's. A naive conversion once stamped the whole document two hours off — including the header every piece of temporal arithmetic is done against
[Goal] · [Tasks]The committed plan, with per-step status: pending, in progress, completed, skipped
[PROPOSED PLAN — NOT COMMITTED]Present only while waiting on the user, and hard-labelled. A proposal that reads like a task list is worse than no proposal at all
[Camera focus]The standing lens, plus a running count of close-up frames used — so the drift is visible to the thing causing it
[Open expectations]Time-anchored ones show a countdown; event-anchored ones show the question the camera is being asked
[Looking for]The find list — still open, or found and where
[Earlier this session]Compacted narrative, with time spans preserved rather than only facts
[Known environment facts]Durable spatial memory. Eight of them by the end of the live session
[Recent observations]Raw captions, newest last, bounded at 24

The order is stability-first. Title, tasks, narrative and environment facts change rarely; raw captions change on every tick. So the things that move least go at the top, and consecutive ticks share the longest possible unchanged prefix — which is what the reasoning provider's prefix cache is able to charge less for. Ordering the document by what a reader would find natural would have cost real money on every frame.

The caption buffer is bounded at 24. On overflow the oldest sixteen are summarised into the narrative by one cheap model call, with time spans preserved rather than only facts, plus any durable environment facts worth promoting. Raw captions are never silently dropped, and the freshest window is never compacted.

04

The trigger engine costs nothing, and it is the only thing that starts a sentence

Wall-clock arithmetic is free; tokens are not. The trigger check runs on every tick and every 20-second poll, costs nothing at all, and the reasoning model is only woken when it returns something — or when a frame arrived, or the user spoke.

TriggerFires when
expectation_duea time-anchored expectation passed its deadline unresolved
stale_taskan in-progress task has gone unmentioned for eight minutes
proposal_pendinga plan was proposed and unanswered for 150 seconds
wanted_founda find-list item was seen and the user has not been told — coalesced into one event for all of them
wanted_stucka search hit its miss or unclear budget. Ask the user rather than failing silently

Each one claims its own state on firing, before anything downstream awaits. An unclaimed trigger fires again on the next tick while the first one is still thinking — the same double-fire lesson the timer system taught in v1.

Unprompted speech is then gated by a 90-second politeness budget, and exactly two things bypass it. [URGENT] — physical risk, or work about to be ruined — and the follow-up window that a user's own question opens.

WHY ANSWERING SOMEONE MUST NOT RESET THE GAP

Speaking used to reset the politeness budget. Asked to find the onions, the assistant replied “I'll point them out as soon as they're in view” — and that reply gagged it for the entire 90-second search. It found them at +27 seconds, logged them silently, and said nothing until it was asked again.

Speaking should make the assistant quieter. Being asked should make it more forthcoming. Those pull in opposite directions, so they are tracked as two separate timestamps.

05

The find list, and the door that was closed

“Find the chicken and the onions” opens a search. Every open item is put to the camera by name, in one block, on every frame — one block rather than one question per item, so the list never competes with the four-watch cap and a fourth search cannot silently stop reaching the camera.

chicken packet: FOUND — middle shelf, behind the milk
onions: NOT VISIBLE — this part of the frame is the door rack

There is no tool that marks an item found. The model can open a search and cancel one; only plain string matching over the caption — zero tokens — sets the found state. That is what stops an inference standing in for an observation. The caption said “several bags of lentils” and the model once upgraded it to “I can see the beans”. The found state now has no model-facing door.

Two smaller decisions here both came from thinking about what happens when something goes wrong. Found and announced are separate flags— one is about the world, the other about speech — so a failed utterance simply re-fires next tick instead of leaving an item found and unspoken forever. And matching is by name, never by question index: the index is positional over a list rebuilt each tick while the vision call runs with the lock released, so eventually it would announce one item's location under another item's name.

06

What one minute of watching costs

In this split the vision call is the only image cost, so its prompt tokens are the per-frame bill — which is what makes the number knowable at all. A tick is roughly 1,350 input and 90 output tokens at the coarse tier; the close-up tier caps the longest side at 1,024 px, which is about a thousand image tokens on its own.

Free tier, 8k/min · 200k/dayDeepInfra, metered
fastest sustainable tickone every ~11 sno per-minute ceiling
a whole day's allowance~139 ticks, totalabout four cents
1,000 ticksnot reachable~$0.26

v1 survives on a free tier because its cadence is slower and most of its turns are text-only. v2 ticks continuously by design — it is a fundamentally heavier vision consumer, and the honest conclusion was that the free tier is not a constraint it can be engineered inside. That is now enforced rather than documented: the wrong provider refuses to start.

The close-up tier gets a nudge, then a backstop. The count of close-up frames is rendered into the document, so the drift is visible to the thing causing it, and a hard cap forces coarse at 120. v1 proved that a fine mode set once is never voluntarily reverted; v2's own live session proved the opposite failure, and the tier was never chosen at all.

07

Fifteen tools, and two that are deliberately missing

ToolWhat it does
propose_plan · commit_plan · discard_planThe approval cycle. A proposal is spoken, never written to tasks, and nothing tracks it until the user agrees
add_wanted · drop_wantedOpens and cancels a search. Neither one can mark an item found — only the camera can do that
update_tasks · mark_taskFor a plan the user is already working through, rather than one being proposed
set_expectation · resolve_expectationA deadline, or a condition put to the camera verbatim on every frame until it resolves
set_vision_focusThe standing lens. Replaced, never appended, so duplicates are impossible by construction
log_environment · retract_environment_factDurable spatial memory, and undoing it. The retraction takes a correction, not just a deletion
web_search · fetch_page · calculateInherited from v1 unchanged, and flagged blocking so they never run on the event loop

The document-mutating tools close over the agent's current in-memory document, so a turn's tool calls and the agent's own writes can never interleave on disk. And retract_environment_fact takes a correction, not just a deletion — the raw captions that produced the wrong inference are still in the buffer and will suggest it again on the very next tick. A hole in the fact list does not block that; a durable “the bag on the pantry shelf is NOT toor dal” does.

start_timerSubsumed by a time-anchored expectation, which also has a resolution path a timer never had — an expectation can be met early, or turn out not to apply
request_camera · request_live_searchv1 tools. The live UI owns the camera now, and chat turns attach the current frame client-side, so there is nothing left for the model to ask for
08

Four designs I rejected

ConsideredRejected because
A VideoLLM-online streaming architectureExamined seriously as a route to real-time. Its pooled per-frame vector is an unconditional compression that cannot preserve a label or a torque figure, and its cache grows without bound so latency degrades over a session. The lesson taken instead: speech latency is not a real ceiling, because synthesis can start before the next frame is processed
A duplicate-watch checkBuilt, measured, removed. The duplicates that actually occurred shared about 19% of their words — “car is safe to work under” versus “the car should be settled firmly on both ramps” — while a threshold low enough to catch those merged genuinely distinct watches differing only in the object. Silently dropping a watch the user is relying on is far worse than carrying a duplicate
A hard cutoff on stale searchesA search past its budget now stays open but stops buying full resolution, and the prompt gets a nudge to close it. Dropping it outright would silently abandon something the user may still be waiting on
A multi-agent orchestratorRejected for v1 and the reasoning holds. The reasoning model orchestrates itself — it reads the document, decides mid-thought whether to call a tool, and routes to the right response. The thinking chain is the orchestration

The duplicate check is the one I would highlight. It was a reasonable idea, it was built, and measuring it is what killed it — in both directions at once. Removing it also removed the pressure it existed for, because form and safety moved into a tool that replaces rather than appends, which makes duplicates impossible by construction instead of by threshold.