Pre-pilot · research prototype · no production clinical writes

Home/Blog

LLM integrationFault injectionPythonSarvam

sarvam-105b returned null. Tripling the token budget didn't help.

Our first live extraction call spent its whole token budget on reasoning. The fix, then a fault-injection run of all three versions of the provider against 14 failure modes, and the open bug it found.

AI Frontier India engineering4 October 202613 min readverified against d73c7e1
Contents
  1. 1.Where extraction sits
  2. 2.What the provider reads
  3. 3.The first live calls
  4. 4.reasoning_effort: null
  5. 5.Three versions, two diffs
  6. 6.Fault injection across versions
  7. 7.The results matrix
  8. 8.What the matrix shows
  9. 9.An open bug: two calls, no audit
  10. 10.What we have not measured
  11. 11.Reproduce it
  12. 12.If you call a reasoning model

Where extraction sits

The ENT Clinic app turns a doctor's dictated examination into a structured record. Two layers call a model. The layer that decides what enters the record is plain code, and only the clinician makes it final.

  1. L2 · model

    Speech to text

    speech_to_text.py · saaras:v3

    Audio in, transcript out.

  2. L3 · model

    Extraction

    extraction.py · sarvam-105b

    Transcript in, candidate findings out. Each one quotes the transcript.

  3. L4 · deterministic

    Clinical model

    validate_candidate()

    Accepts a candidate as a finding or rejects it. No model involved.

  4. L5 · model

    Draft generation

    consultation_note

    Note and outputs drafted from accepted findings.

  5. L6 · clinician

    Review and sign

    signed in the app

    The only place a record becomes real.

Layers L2 to L6 from docs/03_architecture/system/ARCHITECTURE.md. L4 is entos.extraction.validate_candidate. This post is about L3.

L3 has one job: send the transcript to sarvam-105b and return a list of candidate dicts. Each item names a structure, an observation, a value, a side, and the exact words that support it. Whatever L3 returns is still only a candidate.

That decides how L3 must fail. Everything downstream reads an empty list as "the doctor said nothing worth recording." In an ear exam, that looks like a normal exam. So L3 has to keep two outcomes apart. If the model found nothing, L3 returns []. If the model produced nothing usable, it raises SarvamUnavailable. The callers catch exactly that class:

modules/consultation/implementation/workflow.py:492-500 (generate_proposals)
try:
    extract = self.extraction_provider.extract
    try:
        candidates = extract(transcript, schema=self.schema)
    except TypeError:
        candidates = extract(transcript)
except SarvamUnavailable as exc:
    self._audit_ai_outage(encounter, layer="extraction", reason=str(exc))
    raise
The caller's contract. Only SarvamUnavailable is audited as an AI outage. The API route turns it into a 503 and commits the audit row first.

Keep the inner except TypeError in mind. It lets providers that take no schema argument still work. It comes back in section 9.

What the provider reads

Sarvam's chat completions endpoint returns an OpenAI-style envelope. The provider reads four things from it, and every failure in this post is one of them going wrong:

FieldWhat it means for extraction
HTTP statusAnything but 2xx means no answer.
choices[0].finish_reasonstop means the model finished. length means it hit max_tokens and the output is incomplete.
choices[0].message.contentThe answer text. Expected to contain one JSON array.
choices[0].message.reasoning_contentThe model's reasoning. The provider never reads it, but it counts against max_tokens.

The last row is the whole bug. On a reasoning model, max_tokens is one budget shared by reasoning and answer. If the reasoning uses all of it, you get finish_reason: "length" and content: null.

The first live calls

On 12 September 2026 we put a Sarvam key in the development environment. Before then, every extraction the app had shown came from a fixture. The provider at that commit looked like this:

ai/providers.py @ e4582d9
class SarvamClinicalExtractionProvider:
    """...
    Inherits the [...] behavior documented in [the internal test harness client]:
    one retry at 3x the token budget, `reasoning_effort: "low"`, and a raised
    `SarvamUnavailable` rather than a silent empty result if it still fails
    """
    ...
        max_tokens = 800
        payload = {..., "max_tokens": max_tokens, "reasoning_effort": "low"}
        for attempt in range(2):
            resp = requests.post(..., json=payload, timeout=30)
            resp.raise_for_status()
            data = resp.json()
            content = data["choices"][0]["message"]["content"]
            if content:
                last_content = content
                break
            max_tokens *= 3  # documented sarvam-105b reasoning-budget behavior
            payload["max_tokens"] = max_tokens
Excerpt. The retry policy was copied from an internal test harness client, where answers were short. Nobody re-measured it for extraction.

We made five live calls that session, about 6,900 completion tokens in total, and checked each result before making the next. The log describes these attempts:

#reasoning_effortmax_tokensResult
1"low"800finish_reason length, content nullNo answer
2"low"2400120 s client timeout, returned after 40.6 s with 9,003 characters of reasoning_content, content nullNo answer
3"none"HTTP 400, no tokens billedRejected
4null800Structured JSON array, same demo transcriptAnswer
From REPOSITORY-BASELINE.md (12 September 2026) and the module docstring. These are logged fields, not captured responses. We did not log the usage block, so there are no per-call token counts.

Attempt 2 rules out the obvious explanations. It had three times the budget and four times the code's 30 second timeout. It came back in 40.6 seconds, well inside the timeout, so this was not a slow network. The extra budget went to more reasoning, and the answer never started.

reasoning_effort: null

We first tried "none", because other APIs use that string. Sarvam returned HTTP 400 and billed nothing. The chat completions reference on docs.sarvam.ai, checked the same day, types the field as:

reasoning_effort: "low" | "medium" | "high" | null

An explicit null turns reasoning off. In Python that is None in the payload dict, which requests serialises as JSON null. The key is sent on every request. Our log has no test of what the API does when the key is left out, so we don't rely on the server default.

Three versions, two diffs

The provider has had three versions since then. We call them v0, v1 and v2 for the rest of this post. The first diff is the 12 September fix. It landed in the same commit that split providers.py into one file per layer:

v0 ai/providers.py @ e4582d9 → v1 ai/extraction.py @ 21616a1
@@ extract() @@-        max_tokens = 800         payload: dict[str, Any] = {             "model": self.model_id,             "messages": [{"role": "user", "content": self._PROMPT.format(transcript=transcript)}],-            "max_tokens": max_tokens,-            "reasoning_effort": "low",+            "max_tokens": 800,+            "reasoning_effort": None,  # disables reasoning; see module docstring         }@@ retry loop @@             if content:                 last_content = content                 break-            max_tokens *= 3  # documented sarvam-105b reasoning-budget behavior-            payload["max_tokens"] = max_tokens
Excerpt. The commit also moved the class and renamed the loop variable. max_tokens stayed at 800.

The second diff came on 3 October, when the app began extracting a whole session in one call:

v1 21616a1 → v2 modules/ai/implementation/extraction.py @ d3e99f9
-            "max_tokens": 800,+            # A whole session goes in at once, so the array can be long.+            "max_tokens": 2048, ...-            resp = requests.post(..., timeout=30)-            resp.raise_for_status()+            try:+                resp = requests.post(..., timeout=60)+            except requests.RequestException as exc:+                raise SarvamUnavailable(f"extraction request failed: {type(exc).__name__}") from exc+            if not resp.ok:+                raise SarvamUnavailable(f"extraction returned HTTP {resp.status_code}: ...")             data = resp.json()-            content = data["choices"][0]["message"]["content"]+            choice = data["choices"][0]+            if choice.get("finish_reason") == "length":+                # A cut-off array would parse as fewer findings than were said.+                raise SarvamUnavailable("extraction hit the token limit before finishing")+            content = choice["message"]["content"] ...-        return json.loads(match.group(0))+        try:+            return json.loads(match.group(0))+        except json.JSONDecodeError as exc:+            raise SarvamUnavailable(f"the model's JSON array did not parse: {exc}") from exc
Excerpt from 3 October 2026.

The comment on the length guard makes a specific claim: without it, a cut-off array would parse as fewer findings. We wanted to check that claim, not repeat it.

Fault injection across versions

The question is simple. Given the same bad response, what does each version actually do? To answer it, we run the shipped source of all three versions, not a rewrite of them. The harness loads v0 and v1 straight from git and imports v2 from the working tree:

scripts/fault_injection_extraction.py
def load_v0():
    src = git_show("e4582d9", "domains/ent-clinic-os/08_implementation/ai/providers.py")
    mod = types.ModuleType("fi_v0")
    sys.modules["fi_v0"] = mod
    exec(compile(src, "e4582d9:ai/providers.py", "exec"), mod.__dict__)
    return mod.SarvamClinicalExtractionProvider, mod.SarvamUnavailable
Excerpt. Each version gets its own SarvamUnavailable class, so an exception counts as wrapped only if it is that version's own type.

Each case is a script of HTTP responses. requests.post is replaced for the whole run and the API key is a dummy, so nothing can reach Sarvam. The fake returns real requests.Response objects. That way raise_for_status(), .ok and .json() behave exactly as they do in production, not as a stub would:

scripts/fault_injection_extraction.py
def fake_post(url, headers=None, json=None, timeout=None, **_):
    sent.append({"max_tokens": json.get("max_tokens"),
                 "reasoning_effort": json.get("reasoning_effort"),
                 "timeout_s": timeout})
    spec = script[min(len(sent) - 1, len(script) - 1)]
    if spec.get("raise") == "timeout":
        raise requests.Timeout("read timed out")
    return response(spec)  # a real requests.models.Response

The results matrix

Injected responsev0 e4582d9v1 21616a1v2 current
Complete array, finish_reason stop2 items×12 items×12 items×1
Empty array: the model found nothing0 items×10 items×10 items×1
Array wrapped in prose and a markdown fence2 items×12 items×12 items×1
Prose only, no arrayUnavailable×1Unavailable×1Unavailable×1
Reasoning used the whole budget: length, content null, every callUnavailable×2Unavailable×2Unavailable×1
Null content once (length), then a complete answer2 items×22 items×2Unavailable×1
Null content with finish_reason stop, every callUnavailable×2Unavailable×2Unavailable×2
Cut off by the token limit inside the second itemUnavailable×1Unavailable×1Unavailable×1
Array complete, then cut off in trailing prose2 items×12 items×1Unavailable×1
Constructed: cut off just after a nested list inside a qualifierJSONDecodeError×1JSONDecodeError×1Unavailable×1
HTTP 400 from the APIHTTPError×1HTTPError×1Unavailable×1
HTTP 503 from the APIHTTPError×1HTTPError×1Unavailable×1
Client timeout before any responseTimeout×1Timeout×1Unavailable×1
Malformed envelope: choices is nullTypeError×1TypeError×1TypeError×1
Generated from docs/09_verification/fault-injection-extraction.json. ×N is the number of HTTP requests the version made. Green returned a list. Amber raised that version's own SarvamUnavailable. Red let a different exception type escape, unwrapped.

The requests each version sent for the reasoning-exhaustion case:

VersionRequest 1Request 2
v0800 · "low" · 30 s2400 · "low" · 30 s
v1800 · null · 30 s800 · null · 30 s
v22048 · null · 60 snone
max_tokens · reasoning_effort · client timeout, as recorded by the fake.

What the matrix shows

1. A cut-off array never became a shorter list

Two cases cut the array itself: once inside an item, once just after a nested list. Every version failed on both. None returned fewer items. The reason is the parser. Every version pulls the array out with a greedy regex:

match = re.search(r"\[.*\]", last_content, re.DOTALL)

A cut-off array has no closing bracket, so the match fails and the provider raises "could not find a JSON array." That is safe, but only by accident. When the cut falls right after an inner ], the regex matches a broken prefix and json.loads raises. In v0 and v1 that error was a raw JSONDecodeError.

So the comment on the v2 guard describes a failure we could not produce. The guard is still worth having, for a different reason than the one it states. It turns every cut-off into one typed error with one message, whatever the cut-off text happens to look like.

2. v0 and v1 let three exception types past the audit

HTTP 400, HTTP 503 and timeouts escaped v0 and v1 as raw HTTPError and Timeout. The nested truncation escaped as JSONDecodeError. The caller in section 1 catches only SarvamUnavailable. Those failures would have skipped the outage audit row and the 503 mapping. v2 wraps all of them.

3. v2 is stricter, and that costs something

Two rows went from green to amber. In "null content once, then a complete answer," v0 and v1 recovered on their second request. v2 never makes one, because the length check runs before the null check, and its retry loop now only fires on null content when finish_reason is not length. In "array complete, then cut off," v2 refuses a usable array because the trailing prose was truncated.

We accept both. A false failure is visible and costs the doctor a retry. A wrong record costs more. But this is a trade-off we chose, and it should be stated as one.

An open bug: two calls, no audit

The last row is red in every version. If the envelope is malformed, for example {"choices": null}, then data["choices"][0] raises TypeError inside extract(). Now go back to the caller in section 1. It reads any TypeError as "this provider takes no schema argument" and calls extract() again.

We ran this through the real ClinicalWorkflowService.generate_proposals with a patched transport. The result: 2 HTTP requests, a raw TypeError, and 0 outage rows in audit_events. Two tests now pin this:

tests/integration/test_extraction_envelope.py
def test_today_a_malformed_envelope_costs_two_calls_and_no_audit_row(run):
    error, http_calls, outages = run()
    assert isinstance(error, TypeError)
    assert http_calls == 2
    assert outages == 0


@pytest.mark.xfail(strict=True, reason="known gap: envelope errors are not wrapped in SarvamUnavailable")
def test_a_malformed_envelope_should_be_one_call_and_an_audited_outage(run):
    error, http_calls, outages = run()
    assert http_calls == 1
    assert outages == 1
The first test records today's behavior. The second states the behavior we want, and is a strict xfail until the fix lands. When the bug is fixed, the xfail starts passing and the suite fails, which forces the tests to be updated.

There are two separate defects. The provider should wrap envelope errors in SarvamUnavailable. The caller should not use TypeError to detect a signature. Checking for the parameter, or catching only the error raised at call binding, would do it. The same caller pattern appears in analyze_findings at workflow.py:422. We tested only generate_proposals. Neither defect is fixed yet.

What we have not measured

ClaimStatusEvidence
sarvam-105b returns structured output for our demo transcript with reasoning offTested live12 Sep 2026, one transcript
Each version behaves as in the matrix for these 14 inputsTested offlined73c7e1, synthetic responses
Malformed envelope causes a second call and no audit rowTested offlineStrict xfail pinned
How often Sarvam actually returns each faultNot measuredNeeds logged production traffic
The extracted findings are correctNot measuredHarness NOT_RUN, no clinician-labeled set
Reasoning off does not hurt extraction qualityNot measuredNo comparison run

The matrix tells you what the code does with a given response. It tells you nothing about how often Sarvam sends that response. And one answer that parses is not an accuracy number. The evaluation harness in modules/evaluation still reports NOT_RUN, because nobody has built a set of transcripts with findings a clinician confirmed.

Also, once we had a live answer, our own L4 rejected every candidate in it. The reason had nothing to do with medicine. That is the next post.

Reproduce it

$ python scripts/fault_injection_extraction.py
complete           2 items [1 call] | 2 items [1 call] | 2 items [1 call]
...
no_choices         TypeError (unwrapped) [1 call] | TypeError (unwrapped) [1 call] | TypeError (unwrapped) [1 call]
wrote docs/09_verification/fault-injection-extraction.json

$ python -m pytest -q
250 passed, 2 xfailed
From domains/ent-clinic-os. No API key needed. Nothing leaves the machine.

If you call a reasoning model

  • Read finish_reason before content. On a reasoning model, max_tokens is shared, and null content with length means the reasoning used it up.
  • Turn reasoning off for fixed-shape extraction, using the exact value the provider documents. "none" and null are different requests.
  • Don't carry a retry policy across tasks. Re-measure it on the task you ship.
  • Wrap every failure your caller can't act on in one exception type. Transport, envelope and parse errors included. Then test each path with real Response objects.
  • Never detect a function's signature with except TypeError around the call. It also catches every TypeError raised inside it.
  • When you harden error handling, run the old and new code on the same faults. The diff in behavior is the honest changelog.
  • Log the usage block on every call. We didn't, and it shows above.