ENGINEERING CASE STUDY · v2 PROTOTYPE · ONE LIVE RUN
Chitragupta
A live vision assistant, without the GPUs, the dataset or the trained model the research assumes.
WORLD-DOCUMENT ARCHITECTURE · QWEN3-VL-30B ON DEEPINFRA + DEEPSEEK V4-FLASH
A hands-free camera assistant for hands-on work — cooking, repairs, a supermarket aisle. It watches through a phone camera, keeps a written record of everything it has seen and decided, and speaks only when speaking is worth it. The user is listening, not reading, which drives almost every design decision below.
THE PROBLEM
In the literature, a live camera assistant is a training problem. I had no GPUs, no dataset and nothing to fine-tune — so it had to become an architecture problem instead.
This began as a master's seminar. The paper I presented was VideoLLM-online, and it is the right place to start because it asks the harder of the two questions. Not what is in this frame — captioning solved that — but when a model watching a live stream should speak at all. An assistant that answers the first question well and the second one badly is not a weak assistant; it is an unusable one, because it narrates.
The paper answers it with a trained streaming head. I had a laptop and free-tier API keys. So the question I actually had to answer was whether that behaviour survives when every mechanism that produces it is out of reach — and this is the shape of the answer.
for a question to reach the wire while a tick is mid-reasoning. It was 2.00 s when one lock covered the whole tick
to decide whether the user is owed something. The trigger engine is arithmetic over the document, not a model call
measured vision cost per thousand ticks on DeepInfra. Matching a whole free-tier daily allowance costs about four cents
of real traffic. v2 has been run live exactly once, on 2026-08-10, and everything on this page is written from that export
The shape of it, and why it has that shape
Two models that cannot see what the other sees, a document between them, and a trigger engine that costs nothing. The reasoning model owns every decision and every tool call and has never seen a pixel in its life; the vision model is looking straight at the answer and knows nothing about the conversation.
There are two established ways to build this, and both assume a budget
Train a streaming model. VideoLLM-online runs each frame through a CLIP tower, pools it to a single vector, and appends it to a context that grows as the video plays, with a trained head firing at every step to decide whether this is a moment worth speaking at. The temporal grounding, the speak-or-stay-quiet decision and the instruction data are all trained artifacts.
Or reach for a realtime multimodal model. Serve an open-weights VLM like Qwen3-VL yourself, or rent a hosted live API such as Gemini's. Perception here is genuinely solved — a pooled CLIP vector cannot preserve a label or a torque figure, and these read both off a phone frame — but you are paying for GPUs to serve it, or for a stream metered by the second whether or not anything happened.
The two differ in almost every respect except the one that mattered to me: both put the intelligence inside a single model, and both assume hardware or a meter. I had a laptop and free-tier API keys.
| Approach | What it needs | What running it costs | What you can change |
|---|---|---|---|
| Train a streaming modelVideoLLM-online | GPUs, and instruction data purpose-built for every task it assists with | Cheap per frame once trained. The training run is the bill | Anything — by retraining. Which means nothing, without the hardware |
| Run a realtime multimodal modelQwen3-VL self-hosted · Gemini Live | GPUs to serve it, or a hosted stream you do not control | Metered by the second of stream, whether or not anything happened | The prompt. When it speaks is the model's judgement, and you cannot inspect it |
| Compose hosted models around a documentthis project | Two API keys and a phone | ~$0.26 per 1,000 ticks — and an unchanged scene never becomes a request at all | Everything that matters is code. The silence policy is arithmetic you can read and test |
The third route: composition instead of capacity
Keep the models hosted and interchangeable, and put the intelligence in the architecture between them. That is not only the cheap option — it is the one where the behaviour everyone cares about stops being a learned weight and becomes code you can read. When a trained head speaks at the wrong moment you retrain it. When arithmetic speaks at the wrong moment you open the file and change a number, and you can write a test for it.
The same seam buys capabilities neither approach has a natural place for, because the reasoning stage is a general tool-calling model rather than a video head:
- Web search at runtime — it can attempt a task nobody prepared data for. The trained approach's answer to an unfamiliar job is a new dataset; this one's is a search and a page fetch, mid-conversation.
- Time-anchored expectations — a deadline stored as wall-clock and checked by subtraction, which costs no inference and survives a restart. Tracking elapsed time from a frame stream is the paper's weakest point, and it does not improve on a sampled feed.
- State that is a file, not a context window — the plan, the environment facts and the search list are on disk. A streaming session's memory dies with the connection; this is still there tomorrow.
- A silence policy you can audit — why it did or did not speak at any moment is arithmetic over that file, so it can be read, tuned and unit-tested — rather than inferred from a model's behaviour.
To be fair to the alternatives: hosted live APIs do offer function calling inside a session, so tool use is not unique here. The difference worth defending is that this state outlives the session — it is a file, so it survives the connection dropping, the server restarting, and the user closing the app and coming back after lunch.
Every job the paper does with training has to move somewhere cheaper. Temporal continuity moves into text. Purpose-built instruction data becomes a web search. Elapsed time becomes wall-clock subtraction. Those are all fine. But the frame-by-frame decision to speak — the paper's whole contribution — has nowhere obvious to go except into the prompt, which is exactly where it fails.
The solution
The first version put that decision in the prompt and asked it on every single frame: what do I say about this? That question has an answer every time, which is why it narrated. Silence had to be bolted on afterwards — a [SILENT] protocol, a regex to catch the model narrating its own silence, and a list of near-variants to strip. Three mechanisms all fighting the framing of the prompt they were attached to.
v2 stops asking it. A tick's only job is to make a written record match what is now true — which is a question with a checkable answer, and one a model is good at. Whether the user is owed a sentence is then decided separately, from that record, by arithmetic that costs no tokens and involves no model. Most ticks correctly produce no speech, and not because silence was enforced: producing speech was never the objective being scored.
That is the substitution for the trained head, and everything else falls out of it. State survives restarts because it is a file rather than a conversation. The cost model works because the expensive question is asked rarely and the free one continuously. And the failure that prompted the rewrite — reading a label into the record and then saying nothing about it — became fixable, because the two jobs it was failing to balance are now two separate calls.
THE THESIS
If you cannot train the decision to speak, stop asking a model to make it on every frame. The document is what is true; speech is a side-effect of it.
What one tick actually does
The diagram above is a still picture of a loop. This is the loop itself: everything between a frame arriving and either a sentence or silence, in the order it happens. The two columns worth reading down are what each phase costs and whether the document is locked while it runs.
The frame plus a standing brief go to a model that reads scenes and writes prose. It never reasons and never decides.
A trigger claims its own state before anything awaits, or it fires again on the next tick while the first one is still thinking.
Scored on one thing: is the document now accurate? Its prose is thrown away unread.
Arithmetic. When nothing happened, the next phase never runs at all — which is what makes an idle tick cost exactly one reasoning call.
A small, separate prompt: no tools, no system brief, no full document. One question, and the user's own words.
90 seconds between unprompted remarks, bypassed by [URGENT] and by the follow-up window a user's question opens.
Splitting phase 3b from 3d fixed a failure reported three times, where the model read a label into the document and then said nothing about it. One call was being scored on two objectives that pull against each other — keep the document accurate, which rewards quiet bookkeeping, and decide whether to speak, which rewards noticing the user. Silence kept winning, because it was also the stated default for the other job.
The lock column is the other thing to notice. The document is a single JSON file, so load → mutate → save has to be atomic and there is exactly one lock — but it is held only across the three millisecond-long phases, and never across a model call. It used to cover the whole tick, which meant a person asking a question queued behind two model round trips and any web search they made.
The price of releasing it is that a turn now reasons about a document that may have moved by the time it writes — which is why every window reloads from disk, and why the document-mutating tools degrade to a harmless "no match" string rather than raising when the thing they name is gone.
Losing a tick's bookkeeping to a race is recoverable — the next frame re-derives it. Making the user wait is not.
Measured at the moment a question arrives mid-reasoning, that took time-to-wire from 2.00 s to 0.01 s. How it works has the timeline diagram and the rest of the concurrency argument.
Aiming a camera you cannot look through
The reasoning model cannot see, so it aims the camera in writing, two ways: a standing lens describing what the user is physically doing, and discrete watches — conditions put to the camera verbatim as questions on every frame until they resolve.
set_vision_focus(brief = "User is dicing onions on a board at the counter")
Q1: FOUND — <exactly where, plus any label text> Q1: NOT VISIBLE — <what is in that part of the frame instead> Q1: UNCLEAR — <what you can make out, and what is blocking an answer>
Three rules here each cost a debugging session. The first is the one I would defend hardest: a brief must ask for observations, never judgement. A model asked “is the grip safe?” reaches for reassurance; asked “are the fingertips curled back or extended flat?” it returns a fact. A wrong reassurance is the most damaging thing this stage can produce, so the prompt says outright that safe, correct, proper and fine are not words the camera is allowed to use.
The second: the reasoning model writes only the activity. When it wrote the whole vision block it wrote checklists — “whether a drain pan is directly underneath the filter” — which enumerates one imagined arrangement, so a different but perfectly fine setup read back as a list of absent items. It cannot see the user's kitchen, so it does not get to describe it. The grip and posture wording is standard text attached to every frame.
The third: NOT VISIBLE is a real answer, and it is asked first, before the description. Without an explicit question the brief arrived as a soft “these are relevant” nudge and nothing sharp came back. Observed live: the user wanted black-eyed beans, the caption said “several bags of lentils”, and the reasoning model — with no answer to read — upgraded that into “I can see the beans”. An inference stood in for an observation because no question was posed.
Plans are proposed, not written
The model writes two categorically different things, and they need different rules.
Observations — a caption, an environment fact, a resolved expectation — are reports of what it saw. They write silently and immediately, because gating them on approval would turn every tick into a permission prompt and destroy the point of a hands-free assistant.
A plan is a decision about how the user spends the next hour, made from a photograph and a web search. Once it lands in the task list it is re-injected into every later prompt as settled fact — so the model reads its own guess back as memory and holds the user to it.
propose_plan → a proposal, NOT tasks. Nothing tracks it, no expectations.
Must be said out loud in the same reply — the user cannot see it.
commit_plan → promotes it to real tasks. Fires on assent, including the
implicit kind: “yes”, “go on”, or visibly starting step one.
discard_plan → drops it.An unanswered proposal re-raises itself every 150 seconds, because this is the one silence that costs the user something rather than sparing them: the assistant is blocked on them, they have no idea it is waiting, and nothing is being tracked in the meantime. And a tick may only commit on visibly starting step one — a frame cannot tell you someone said yes.
What one live session actually showed
v2 has been run against real traffic once — an eighteen-minute chicken-curry session on 2026-08-10, exported as fifty-six entries. Six things it got right and three it did not, from the same eighteen minutes. Quoted from that export, via this project's decision log.
“I'm still waiting to hear if that style works for you”
Said unprompted, minutes after proposing a six-step plan. A proposal nobody answered re-raises itself, because the assistant is blocked on the user and the user has no idea it is waiting.
“yes this style works for me”
The user's assent, on which commit_plan fired together with update_tasks and set_vision_focus. Until that line, the plan existed only as a proposal and nothing was tracking it.
“my view is completely black (looks like a lens cap)… carry me with you and the moment the view clears I'll look”
A dead camera reported as a dead camera. To a frame-difference comparison a black frame is a perfectly still scene, so this needed a separate liveness test to be sayable at all.
“⚠️ Slow down—path is tight by the door and shelving”
[URGENT], fired off a motion-blurred frame while the user was walking. It bypasses the 90-second politeness gate and does nothing else — one flag, one consequence.
“BREASTS — keep short to avoid drying out”
A mid-plan correction — breasts, not thighs — propagated into two task notes with the reasoning attached, so it is read back at the step where it changes what to do.
“capture: user turn — scene unchanged, reusing the last caption”
An idle question answered without paying for a fresh caption. The server tells the model this outright, rather than letting it conclude it has gone blind.
“Let me also make a note of what's on the menu so it stays consistent:”
The end of the turn. A dangling reply that the deterministic repair should have caught, but the guard requires a plan tool in the results and this turn's only tool was log_environment. Still open.
“429 — Rate limit reached for model qwen/qwen3.6-27b … TPD: Limit 200000, Used 199109”
How the session ended: on a provider v2 should never have been able to reach. Every config file on disk said otherwise, which is exactly why it is now checked at startup instead of documented.
“every caption in the session is (coarse)”
The close-up tier was never once chosen live, including while reading a patent binder and searching a fridge. It passes its harness; the model does not reach for it. Unresolved.
The export is the highest-value artifact in this project. It carries every vision prompt and answer, silent ticks included, plus the document as it stood — and reading it top to bottom is how the contradictions were found. No single turn shows them; only the sequence does.
Every architectural decision traces back to a number
The vision model is the only thing here that costs real money, and in this split it does nothing but look at pictures — so its prompt tokens are the per-frame bill. Image cost scales with resolution, not file size: compressing the JPEG harder buys nothing at all, which took measurement to establish rather than assumption. Quality is fixed and is not a lever.
So the close-up tier is opt-in, and the resolution decision has to reach the browser before the next capture — resolution discarded in the client can never be recovered on the server, and nothing can upscale a label back. The frame that causes an upgrade is therefore itself coarse.
A v2 tick is roughly 1,440 tokens against the free tier's 8,000-per-minute cap — one tick every eleven seconds, against a four-second interval — and its 200,000-per-day cap works out to about 139 ticks total, per day. v2 ticks continuously by design. It is a fundamentally heavier vision consumer than v1 and cannot run there at all.
It ran there anyway, and the session died eighteen minutes in on a rate limit. Nothing was wrong on disk — every config file said the right provider, so there was no artifact anyone could inspect and find a mistake in. That is the actual defect: the system had a hard requirement and no point at which it compared that requirement against reality.
The default nobody sets must be the one that works. It defaulted to the one configuration that cannot.
Three changes came out of it, and the middle one is the reusable part. Which provider gets the pixels was not a question any code could have answered: the class name, the mode string and the module name all fail to tell you, because the hybrid backend extends the other one and replaces a client its parent constructed. Backends now declare it. And v2 refuses to start on the wrong one, unless an environment variable says the choice was deliberate — arriving there is the bug; choosing it for a one-off comparison is legitimate.
Initialized live agent | backend mode: deepinfra | VISION ON: deepinfra
| reasoning: deepseek-v4-flashEvery character in that check is ASCII. The first version used arrows and section marks, and printing it on a Windows console raised a UnicodeEncodeError — a diagnostic that crashed while reporting the problem it existed to explain.
Every trained mechanism, and where it had to move
This is the substitution table the whole project is: six things VideoLLM-online does with training, and what each one became once training was off the table. Read against the paper, the substitutions are the argument.
Worth saying plainly: I went back and examined the streaming architecture again while building v2, as a route to real-time, and rejected it on the merits rather than on budget. Two of the rows below would still be the right call with a rack of GPUs behind them.
| The paper does | Why not here | What v2 does instead |
|---|---|---|
| Pools each frame to a single vector | An unconditional compression — it cannot preserve a label, a torque figure, or where fingertips sit relative to a blade | A query-conditioned caption: the brief tells the vision model what matters before it looks, which is why read mode can transcribe a packet at all |
| Grows a KV cache of frame tokens | Decoding is memory-bandwidth bound, so latency degrades over a long session | No image tokens ever accumulate. Each frame is captioned independently and only text persists, bounded at 24 captions with span-preserving compaction behind it |
| Trains an EOS head to decide when to speak | No training budget | A zero-token arithmetic engine over the document decides *whether* to ask, and a separate small prompt answers *what* to say |
| Inference on every frame | A continuous tick loop is the heaviest thing in the system | A perceptual diff gate in the browser — an unchanged scene never becomes a request at all |
| Instruction data built for the tasks it assists with | No dataset, and no budget to build one | It searches the web at runtime, so it can attempt a task nobody prepared data for |
| Tracking elapsed time from the frame stream | The paper's weakest point, and it does not improve on a sampled feed | A time-anchored expectation: a stored deadline checked by subtraction, which also carries a resolution path a timer never had |
The genuine lesson taken from it is the one about latency: speech latency is not a real ceiling. Streaming the reply to synthesis lets the assistant talk while the next frame is already being processed. That is designed and not yet built, and it is the next thing.
Go deeper
The deployed v2 instance. Two things to know before you click: open it on a phone — it wants a camera, a microphone and a secure context, and a laptop webcam pointed at your face is not the thing it was built for. And it is a free-tier host that sleeps after about fifteen minutes, so the first request can take the better part of a minute to wake it before anything happens.
The document itself, section by section; what one tick costs and where; the fifteen tools; and the four designs rejected along the way.
Ten failures with the symptom, the root cause, and the rule each one produced — plus the rules v1 left behind and where they still bind.
Early, and specific about it
- Propose-then-commit: a six-step plan held unwritten across several minutes and many ticks, raised unprompted, committed on assent
- The plan read aloud in the same reply that committed it — the user never had to ask what was in the document
- A dead camera reported honestly and usefully, rather than as “nothing has changed”
- [URGENT] reaching the user immediately, off a motion-blurred frame in a tight passage
- Compaction preserving a time span — one narrative entry covering 13:24:56 to 13:36:58, not just the facts inside it
- Eight durable environment facts by the end, including the fridge shelf the chicken came off
- One eighteen-minute session is one session. Nothing on this page is a distribution
- The close-up tier has never been chosen by the model in live traffic, including while reading small print
- Flat frames are detected and still captioned at full price — five vision calls went on describing a blackout
- Streaming speech (/v2/chat/stream) is designed, with the constraint written down, and not built
- Nothing throttles a twenty-minute simmer, and the diff gate saves nothing at all while the user is walking
- No automated test suite. Verification is a folder of ad-hoc harnesses plus reading the export, and harnesses cannot tell you whether the model behaves
v1 still runs, on its own routes, and is not being developed. It is kept because it is the control: it is the version that narrated, and the argument for everything above is that it stopped.
What this demonstrates
The one I would point at first: I once spent a session unable to explain why the assistant lost its memory — so I shipped instrumentation instead of a fix, because “the model ignored the task list” and “the task list never reached the prompt” look identical from a transcript and need opposite repairs. The provider bug above is the same shape, one version later, and this time the answer was to make the system check its own requirement out loud at startup.
Where the async is actually load-bearing
A framework name in a list proves nothing. This one runs a single worker with a camera firing on an interval and a person able to interrupt it at any moment, so concurrency is a correctness problem, not a résumé line.