TECHDRISHTI RESEARCH INFRA · PROVIDER EVALUATION
Sarvam vs DeepSeek, side by side
The research agent behind the RESEARCH stage (see the case study) can run on either model. 37 live calls per provider, 74 total, against the same fixed question set pulled from a real Stage 1 pipeline output: identity lookups (ambiguous and unambiguous entities) and context/gap questions, grouped by source article. This page is the evidence behind choosing Sarvam for that stage, not just an assertion of it.
Speed, iterations, length
lower is faster
lower means fewer round-trips
lower is more concise
not scored either way, shown for context
Answer quality, judged independently
Scored by an external GPT judge, not by this pipeline and not by me, on the 4 context/gap questions where quality is genuinely hard to call by eye (the 33 identity lookups weren't scored this way, most of them are plain factual overviews with an obvious right answer). Included as a transparent second opinion the quantitative metrics above can't give on their own.
4 of 37 cases judged
DeepSeek wins every judged case outright, by margin.
Click a question to see the actual answers and what the judge flagged.
Hits every sub-topic asked (compilation pipeline, memory, ISAs, compute units, overheads) and is the best-organized of the two. But oversimplifies backend targets (claims ROCm compiles to HSAIL, understates Intel Arc's real pipeline) and overstates portability (“single binary runs on any accelerator”).
Goes to implementation depth Sarvam never reaches: StableHLO, the zml.Platform abstraction, PJRT plugin layer, DMA allocator, zero-copy transfers. Explains the actual lowering chain (StableHLO -> MLIR dialects -> vendor dialects -> native kernels) instead of naming ISAs. Self-flags one claim as [UNVERIFIED], good practice. Slightly verbose, ~25% could be trimmed.
Covers all three asked angles (business risk, performance-parity risk, competitive strategy) in a focused, readable executive-summary style. Stays fairly high-level; a few claims asserted without much nuance (e.g. “changing a flag recompiles the model”).
Surfaces business risks Sarvam omits or only hints at: unclear monetization, contributor-pool limits, dependency on OpenXLA, the CUDA ecosystem moat. Explains why each risk matters rather than just listing it. More verbose, with some repetition around ecosystem advantages.
Cleanly separates terms / implications / differences and covers every key fact (investment, ownership %, co-selling, Xeon integration). The equity percentages (8.2%, 9%, 6.8%) are oddly specific without adding analytical value, and it explains what happened more than why it matters.
Frames the relationship as an evolution: investor relationship -> board representation -> acquisition talks -> failed acquisition -> strategic alliance, which directly answers the “how does it differ from before” half of the question. Adds useful context (a boardroom recusal, a SoftBank deployment, a later funding round), though the later Series F detail runs past what was asked.
Correctly summarizes the acquisition-status chronology (talks stalled, partnership replaced acquisition). Its IPO/independence discussion leans on generic corporate-finance factors (interest rates, antitrust, timing) rather than SambaNova-specific evidence, and honestly flags this with [UNVERIFIED: IPO not found in retrieved sources].
Same acquisition-status chronology, sourced with a clearer timeline (a rumor, a wire-service report, the funding round, the strategic collaboration). Grounds the independence/IPO analysis in company-specific signals: Series E runway, record bookings, manufacturing expansion, customer traction, not just generic IPO theory. Also appropriately flags the unverified parts.
Why Sarvam anyway
DeepSeek is genuinely better on the questions that need real depth, an independent judge scored it higher on all 4 judged cases, and it goes further into implementation detail (one answer, on ZML's compiler stack, named the exact lowering chain from StableHLO through vendor dialects to native kernels; Sarvam's equivalent answer stayed a level higher, correct but less deep). That's a real, honest edge, not explained away here.
But it takes about 65% longer per call (33.4s vs 20.2s), needs more search iterations to get there (6.1 vs 4.5), and returns roughly double the answer length (512 vs 264 tokens) for that gain, on a stage that fires several times per article, every day, unattended. Combined with Sarvam running 5–9× cheaper per raw token (see the pricing table in the case study), the tradeoff didn't clear the bar for this pipeline specifically: DeepSeek's extra depth is real, but it's not the kind of gap this stage's job actually needs to close, and Sarvam is, in the external judge's own words, “the stronger choice if the goal is a tight executive summary,” which is exactly what feeds the writer downstream.