Back to portfolio

ENGINEERING CASE STUDY · v2 PROTOTYPE · ONE LIVE RUN

Chitragupta

A live vision assistant, without the GPUs, the dataset or the trained model the research assumes.

WORLD-DOCUMENT ARCHITECTURE · QWEN3-VL-30B ON DEEPINFRA + DEEPSEEK V4-FLASH

A hands-free camera assistant for hands-on work — cooking, repairs, a supermarket aisle. It watches through a phone camera, keeps a written record of everything it has seen and decided, and speaks only when speaking is worth it. The user is listening, not reading, which drives almost every design decision below.

THE PROBLEM

In the literature, a live camera assistant is a training problem. I had no GPUs, no dataset and nothing to fine-tune — so it had to become an architecture problem instead.

This began as a master's seminar. The paper I presented was VideoLLM-online, and it is the right place to start because it asks the harder of the two questions. Not what is in this frame — captioning solved that — but when a model watching a live stream should speak at all. An assistant that answers the first question well and the second one badly is not a weak assistant; it is an unusable one, because it narrates.

The paper answers it with a trained streaming head. I had a laptop and free-tier API keys. So the question I actually had to answer was whether that behaviour survives when every mechanism that produces it is out of reach — and this is the shape of the answer.

0.01s

for a question to reach the wire while a tick is mid-reasoning. It was 2.00 s when one lock covered the whole tick

0 tokens

to decide whether the user is owed something. The trigger engine is arithmetic over the document, not a model call

$0.26/1k

measured vision cost per thousand ticks on DeepInfra. Matching a whole free-tier daily allowance costs about four cents

18 min

of real traffic. v2 has been run live exactly once, on 2026-08-10, and everything on this page is written from that export

01

The shape of it, and why it has that shape

Two models that cannot see what the other sees, a document between them, and a trigger engine that costs nothing. The reasoning model owns every decision and every tool call and has never seen a pixel in its life; the vision model is looking straight at the answer and knows nothing about the conversation.

PHONE CAMERAa frame, on an intervalunchanged, or flat —never leaves the browserDEEPINFRAQwen3-VL-30B-A3Bnever reasons · never decidesa plain-text captionthe pixels stop here — nothing below this line has ever seen an imageTHE WORLD DOCUMENTprimary state — it survives restarts, and it is what every prompt is built fromgoal · tasks · proposed plan · expectations · find listnarrative · environment facts · recent captionsreads · writesreads onlySTAGE 1 · BOOKKEEPINGDeepSeek v4-flash15 tools · text only, never a pixelits prose is discardedTRIGGERS0 tokenspure arithmetic over the documentthe only thing that starts a sentenceScored on one thing: is the documentnow accurate? It is never asked to weighthat against whether to speak.an event — or nothingSTAGE 2 · THE SPEECH DECISION“does the user need to hearsomething?”no tools · no document · one questionpoliteness gate · 90 s[URGENT] bypasses itSPOKEN ALOUDon-device speech synthesisAn idle tick costs one vision calland one reasoning call. Stage 2 isskipped entirely when nothing happened.
Read the two arrows out of the document. One goes to a model, which reads and writes and costs a call. The other goes to arithmetic, which only reads and costs nothing — and it is the arithmetic, not the model, that decides a sentence is owed. Inverting those two is the entire difference between this and the version before it.

There are two established ways to build this, and both assume a budget

Train a streaming model. VideoLLM-online runs each frame through a CLIP tower, pools it to a single vector, and appends it to a context that grows as the video plays, with a trained head firing at every step to decide whether this is a moment worth speaking at. The temporal grounding, the speak-or-stay-quiet decision and the instruction data are all trained artifacts.

Or reach for a realtime multimodal model. Serve an open-weights VLM like Qwen3-VL yourself, or rent a hosted live API such as Gemini's. Perception here is genuinely solved — a pooled CLIP vector cannot preserve a label or a torque figure, and these read both off a phone frame — but you are paying for GPUs to serve it, or for a stream metered by the second whether or not anything happened.

The two differ in almost every respect except the one that mattered to me: both put the intelligence inside a single model, and both assume hardware or a meter. I had a laptop and free-tier API keys.

ApproachWhat it needsWhat running it costsWhat you can change
Train a streaming modelVideoLLM-onlineGPUs, and instruction data purpose-built for every task it assists withCheap per frame once trained. The training run is the billAnything — by retraining. Which means nothing, without the hardware
Run a realtime multimodal modelQwen3-VL self-hosted · Gemini LiveGPUs to serve it, or a hosted stream you do not controlMetered by the second of stream, whether or not anything happenedThe prompt. When it speaks is the model's judgement, and you cannot inspect it
Compose hosted models around a documentthis projectTwo API keys and a phone~$0.26 per 1,000 ticks — and an unchanged scene never becomes a request at allEverything that matters is code. The silence policy is arithmetic you can read and test

The third route: composition instead of capacity

Keep the models hosted and interchangeable, and put the intelligence in the architecture between them. That is not only the cheap option — it is the one where the behaviour everyone cares about stops being a learned weight and becomes code you can read. When a trained head speaks at the wrong moment you retrain it. When arithmetic speaks at the wrong moment you open the file and change a number, and you can write a test for it.

The same seam buys capabilities neither approach has a natural place for, because the reasoning stage is a general tool-calling model rather than a video head:

  • Web search at runtime — it can attempt a task nobody prepared data for. The trained approach's answer to an unfamiliar job is a new dataset; this one's is a search and a page fetch, mid-conversation.
  • Time-anchored expectations — a deadline stored as wall-clock and checked by subtraction, which costs no inference and survives a restart. Tracking elapsed time from a frame stream is the paper's weakest point, and it does not improve on a sampled feed.
  • State that is a file, not a context window — the plan, the environment facts and the search list are on disk. A streaming session's memory dies with the connection; this is still there tomorrow.
  • A silence policy you can audit — why it did or did not speak at any moment is arithmetic over that file, so it can be read, tuned and unit-tested — rather than inferred from a model's behaviour.

To be fair to the alternatives: hosted live APIs do offer function calling inside a session, so tool use is not unique here. The difference worth defending is that this state outlives the session — it is a file, so it survives the connection dropping, the server restarting, and the user closing the app and coming back after lunch.

WHICH LEAVES ONE HARD PROBLEM

Every job the paper does with training has to move somewhere cheaper. Temporal continuity moves into text. Purpose-built instruction data becomes a web search. Elapsed time becomes wall-clock subtraction. Those are all fine. But the frame-by-frame decision to speak — the paper's whole contribution — has nowhere obvious to go except into the prompt, which is exactly where it fails.

The solution

The first version put that decision in the prompt and asked it on every single frame: what do I say about this? That question has an answer every time, which is why it narrated. Silence had to be bolted on afterwards — a [SILENT] protocol, a regex to catch the model narrating its own silence, and a list of near-variants to strip. Three mechanisms all fighting the framing of the prompt they were attached to.

v2 stops asking it. A tick's only job is to make a written record match what is now true — which is a question with a checkable answer, and one a model is good at. Whether the user is owed a sentence is then decided separately, from that record, by arithmetic that costs no tokens and involves no model. Most ticks correctly produce no speech, and not because silence was enforced: producing speech was never the objective being scored.

That is the substitution for the trained head, and everything else falls out of it. State survives restarts because it is a file rather than a conversation. The cost model works because the expensive question is asked rarely and the free one continuously. And the failure that prompted the rewrite — reading a label into the record and then saying nothing about it — became fixable, because the two jobs it was failing to balance are now two separate calls.

THE THESIS

If you cannot train the decision to speak, stop asking a model to make it on every frame. The document is what is true; speech is a side-effect of it.

02

What one tick actually does

The diagram above is a still picture of a loop. This is the loop itself: everything between a frame arriving and either a sentence or silence, in the order it happens. The two columns worth reading down are what each phase costs and whether the document is locked while it runs.

1+2
Caption the framevision
~1.5 s · one vision call

The frame plus a standing brief go to a model that reads scenes and writes prose. It never reasons and never decides.

3a
Fold the caption in, claim triggersdocumentlock held
~1 ms

A trigger claims its own state before anything awaits, or it fires again on the next tick while the first one is still thinking.

3b
Stage 1 — bookkeeping, tools onlyreasoning
~2 s · one reasoning call

Scored on one thing: is the document now accurate? Its prose is thrown away unread.

3c
Is the speech question even worth asking?documentlock held
~1 ms

Arithmetic. When nothing happened, the next phase never runs at all — which is what makes an idle tick cost exactly one reasoning call.

3d
Stage 2 — the speech decisionreasoning
~1 s · skipped when idle

A small, separate prompt: no tools, no system brief, no full document. One question, and the user's own words.

3e
Politeness gate, record the utterancedocumentlock held
~1 ms

90 seconds between unprompted remarks, bypassed by [URGENT] and by the follow-up window a user's question opens.

Splitting phase 3b from 3d fixed a failure reported three times, where the model read a label into the document and then said nothing about it. One call was being scored on two objectives that pull against each other — keep the document accurate, which rewards quiet bookkeeping, and decide whether to speak, which rewards noticing the user. Silence kept winning, because it was also the stated default for the other job.

The lock column is the other thing to notice. The document is a single JSON file, so load → mutate → save has to be atomic and there is exactly one lock — but it is held only across the three millisecond-long phases, and never across a model call. It used to cover the whole tick, which meant a person asking a question queued behind two model round trips and any web search they made.

The price of releasing it is that a turn now reasons about a document that may have moved by the time it writes — which is why every window reloads from disk, and why the document-mutating tools degrade to a harmless "no match" string rather than raising when the thing they name is gone.

Losing a tick's bookkeeping to a race is recoverable — the next frame re-derives it. Making the user wait is not.

Measured at the moment a question arrives mid-reasoning, that took time-to-wire from 2.00 s to 0.01 s. How it works has the timeline diagram and the rest of the concurrency argument.

03

Aiming a camera you cannot look through

The reasoning model cannot see, so it aims the camera in writing, two ways: a standing lens describing what the user is physically doing, and discrete watches — conditions put to the camera verbatim as questions on every frame until they resolve.

→ WHAT THE REASONING MODEL IS ALLOWED TO WRITE
set_vision_focus(brief = "User is dicing onions on a board at the counter")
← AND WHAT THE CAMERA IS ASKED, ON EVERY FRAME
Q1: FOUND      — <exactly where, plus any label text>
Q1: NOT VISIBLE — <what is in that part of the frame instead>
Q1: UNCLEAR    — <what you can make out, and what is blocking an answer>

Three rules here each cost a debugging session. The first is the one I would defend hardest: a brief must ask for observations, never judgement. A model asked “is the grip safe?” reaches for reassurance; asked “are the fingertips curled back or extended flat?” it returns a fact. A wrong reassurance is the most damaging thing this stage can produce, so the prompt says outright that safe, correct, proper and fine are not words the camera is allowed to use.

The second: the reasoning model writes only the activity. When it wrote the whole vision block it wrote checklists — “whether a drain pan is directly underneath the filter” — which enumerates one imagined arrangement, so a different but perfectly fine setup read back as a list of absent items. It cannot see the user's kitchen, so it does not get to describe it. The grip and posture wording is standard text attached to every frame.

The third: NOT VISIBLE is a real answer, and it is asked first, before the description. Without an explicit question the brief arrived as a soft “these are relevant” nudge and nothing sharp came back. Observed live: the user wanted black-eyed beans, the caption said “several bags of lentils”, and the reasoning model — with no answer to read — upgraded that into “I can see the beans”. An inference stood in for an observation because no question was posed.

Which is why nothing but the camera can mark a find. The model can open a search and cancel one; there is no tool that sets an item to found. Only plain string matching over the caption does that — zero tokens, and no model-facing door for an inference to walk through. When a find lands, speech is forced: it routes to a prompt that never offers silence, and if the model declines or errors anyway the server speaks a deterministic sentence built from the location the camera wrote. A forced path the model can talk its way out of is not forced.
04

Plans are proposed, not written

The model writes two categorically different things, and they need different rules.

Observations — a caption, an environment fact, a resolved expectation — are reports of what it saw. They write silently and immediately, because gating them on approval would turn every tick into a permission prompt and destroy the point of a hands-free assistant.

A plan is a decision about how the user spends the next hour, made from a photograph and a web search. Once it lands in the task list it is re-injected into every later prompt as settled fact — so the model reads its own guess back as memory and holds the user to it.

propose_plan  → a proposal, NOT tasks. Nothing tracks it, no expectations.
                Must be said out loud in the same reply — the user cannot see it.
commit_plan   → promotes it to real tasks. Fires on assent, including the
                implicit kind: “yes”, “go on”, or visibly starting step one.
discard_plan  → drops it.

An unanswered proposal re-raises itself every 150 seconds, because this is the one silence that costs the user something rather than sparing them: the assistant is blocked on them, they have no idea it is waiting, and nothing is being tracked in the meantime. And a tick may only commit on visibly starting step one — a frame cannot tell you someone said yes.

05

What one live session actually showed

v2 has been run against real traffic once — an eighteen-minute chicken-curry session on 2026-08-10, exported as fifty-six entries. Six things it got right and three it did not, from the same eighteen minutes. Quoted from that export, via this project's decision log.

“I'm still waiting to hear if that style works for you”

Said unprompted, minutes after proposing a six-step plan. A proposal nobody answered re-raises itself, because the assistant is blocked on the user and the user has no idea it is waiting.

“yes this style works for me”

The user's assent, on which commit_plan fired together with update_tasks and set_vision_focus. Until that line, the plan existed only as a proposal and nothing was tracking it.

“my view is completely black (looks like a lens cap)… carry me with you and the moment the view clears I'll look”

A dead camera reported as a dead camera. To a frame-difference comparison a black frame is a perfectly still scene, so this needed a separate liveness test to be sayable at all.

“⚠️ Slow down—path is tight by the door and shelving”

[URGENT], fired off a motion-blurred frame while the user was walking. It bypasses the 90-second politeness gate and does nothing else — one flag, one consequence.

“BREASTS — keep short to avoid drying out”

A mid-plan correction — breasts, not thighs — propagated into two task notes with the reasoning attached, so it is read back at the step where it changes what to do.

“capture: user turn — scene unchanged, reusing the last caption”

An idle question answered without paying for a fresh caption. The server tells the model this outright, rather than letting it conclude it has gone blind.

“Let me also make a note of what's on the menu so it stays consistent:”

The end of the turn. A dangling reply that the deterministic repair should have caught, but the guard requires a plan tool in the results and this turn's only tool was log_environment. Still open.

“429 — Rate limit reached for model qwen/qwen3.6-27b … TPD: Limit 200000, Used 199109”

How the session ended: on a provider v2 should never have been able to reach. Every config file on disk said otherwise, which is exactly why it is now checked at startup instead of documented.

“every caption in the session is (coarse)”

The close-up tier was never once chosen live, including while reading a patent binder and searching a fridge. It passes its harness; the model does not reach for it. Unresolved.

The export is the highest-value artifact in this project. It carries every vision prompt and answer, silent ticks included, plus the document as it stood — and reading it top to bottom is how the contradictions were found. No single turn shows them; only the sequence does.

06

Every architectural decision traces back to a number

The vision model is the only thing here that costs real money, and in this split it does nothing but look at pictures — so its prompt tokens are the per-frame bill. Image cost scales with resolution, not file size: compressing the JPEG harder buys nothing at all, which took measurement to establish rather than assumption. Quality is fixed and is not a lever.

So the close-up tier is opt-in, and the resolution decision has to reach the browser before the next capture — resolution discarded in the client can never be recovered on the server, and nothing can upscale a label back. The frame that causes an upgrade is therefore itself coarse.

THE CONSTRAINT THAT REWROTE THE PROVIDER LAYER

A v2 tick is roughly 1,440 tokens against the free tier's 8,000-per-minute cap — one tick every eleven seconds, against a four-second interval — and its 200,000-per-day cap works out to about 139 ticks total, per day. v2 ticks continuously by design. It is a fundamentally heavier vision consumer than v1 and cannot run there at all.

It ran there anyway, and the session died eighteen minutes in on a rate limit. Nothing was wrong on disk — every config file said the right provider, so there was no artifact anyone could inspect and find a mistake in. That is the actual defect: the system had a hard requirement and no point at which it compared that requirement against reality.

The default nobody sets must be the one that works. It defaulted to the one configuration that cannot.

Three changes came out of it, and the middle one is the reusable part. Which provider gets the pixels was not a question any code could have answered: the class name, the mode string and the module name all fail to tell you, because the hybrid backend extends the other one and replaces a client its parent constructed. Backends now declare it. And v2 refuses to start on the wrong one, unless an environment variable says the choice was deliberate — arriving there is the bug; choosing it for a one-off comparison is legitimate.

Initialized live agent | backend mode: deepinfra | VISION ON: deepinfra
                    | reasoning: deepseek-v4-flash

Every character in that check is ASCII. The first version used arrows and section marks, and printing it on a Windows console raised a UnicodeEncodeError — a diagnostic that crashed while reporting the problem it existed to explain.

07

Every trained mechanism, and where it had to move

This is the substitution table the whole project is: six things VideoLLM-online does with training, and what each one became once training was off the table. Read against the paper, the substitutions are the argument.

Worth saying plainly: I went back and examined the streaming architecture again while building v2, as a route to real-time, and rejected it on the merits rather than on budget. Two of the rows below would still be the right call with a rack of GPUs behind them.

The paper doesWhy not hereWhat v2 does instead
Pools each frame to a single vectorAn unconditional compression — it cannot preserve a label, a torque figure, or where fingertips sit relative to a bladeA query-conditioned caption: the brief tells the vision model what matters before it looks, which is why read mode can transcribe a packet at all
Grows a KV cache of frame tokensDecoding is memory-bandwidth bound, so latency degrades over a long sessionNo image tokens ever accumulate. Each frame is captioned independently and only text persists, bounded at 24 captions with span-preserving compaction behind it
Trains an EOS head to decide when to speakNo training budgetA zero-token arithmetic engine over the document decides *whether* to ask, and a separate small prompt answers *what* to say
Inference on every frameA continuous tick loop is the heaviest thing in the systemA perceptual diff gate in the browser — an unchanged scene never becomes a request at all
Instruction data built for the tasks it assists withNo dataset, and no budget to build oneIt searches the web at runtime, so it can attempt a task nobody prepared data for
Tracking elapsed time from the frame streamThe paper's weakest point, and it does not improve on a sampled feedA time-anchored expectation: a stored deadline checked by subtraction, which also carries a resolution path a timer never had

The genuine lesson taken from it is the one about latency: speech latency is not a real ceiling. Streaming the reply to synthesis lets the assistant talk while the next frame is already being processed. That is designed and not yet built, and it is the next thing.

08

Go deeper

09

Early, and specific about it

CONFIRMED IN THE ONE LIVE SESSION
  • Propose-then-commit: a six-step plan held unwritten across several minutes and many ticks, raised unprompted, committed on assent
  • The plan read aloud in the same reply that committed it — the user never had to ask what was in the document
  • A dead camera reported honestly and usefully, rather than as “nothing has changed”
  • [URGENT] reaching the user immediately, off a motion-blurred frame in a tight passage
  • Compaction preserving a time span — one narrative entry covering 13:24:56 to 13:36:58, not just the facts inside it
  • Eight durable environment facts by the end, including the fridge shelf the chicken came off
NOT PROVEN
  • One eighteen-minute session is one session. Nothing on this page is a distribution
  • The close-up tier has never been chosen by the model in live traffic, including while reading small print
  • Flat frames are detected and still captioned at full price — five vision calls went on describing a blackout
  • Streaming speech (/v2/chat/stream) is designed, with the constraint written down, and not built
  • Nothing throttles a twenty-minute simmer, and the diff gate saves nothing at all while the user is walking
  • No automated test suite. Verification is a folder of ad-hoc harnesses plus reading the export, and harnesses cannot tell you whether the model behaves

v1 still runs, on its own routes, and is not being developed. It is kept because it is the control: it is the version that narrated, and the argument for everything above is that it stopped.

10

What this demonstrates

Multi-model pipeline designState-machine design for LLM agentsConcurrency under a single lockLLM cost engineeringAgent tool-calling & persistent statePrompt engineering from failure dataObservability for non-deterministic systemsHallucination mitigationResearch → production translationPython · FastAPI · asyncio

The one I would point at first: I once spent a session unable to explain why the assistant lost its memory — so I shipped instrumentation instead of a fix, because “the model ignored the task list” and “the task list never reached the prompt” look identical from a transcript and need opposite repairs. The provider bug above is the same shape, one version later, and this time the answer was to make the system check its own requirement out loud at startup.

Where the async is actually load-bearing

A framework name in a list proves nothing. This one runs a single worker with a camera firing on an interval and a person able to interrupt it at any moment, so concurrency is a correctness problem, not a résumé line.

asyncio.LockHeld only inside a write window — reload, mutate, save, release — measured in milliseconds. Every model call happens outside one, including the reasoning that produces the tool calls. It used to cover the whole tick, and a question arriving mid-reasoning waited 2.00 s to reach the wire; it now waits 0.01 s.
reload-on-entryThe price of releasing the lock is that the document may have moved. Every window reloads from disk unconditionally, because a window reusing a document read before a model call silently rolls back whoever wrote in the meantime — and both writes report success.
asyncio.to_threadNetwork tools are flagged blocking=True and dispatched to a thread. Before that they ran on the event loop, so one slow web search stalled every camera tick behind it.