Ollama skips your JSON schema when a thinking model answers without thinking
Here's a thinking process:
-
Analyze User Request:
- Input: A raw, undifferentiated wall of text from an article titled "Ollama skips your JSON schema when a thinking model answers without thinking"
- Task: Reformat into clean, readable Markdown for a technical audience
- Rules:
- Break into short, well-structured paragraphs separated by blank lines
- Add Markdown section headings (## Heading) where topic clearly shifts
- Use bullet/numbered lists for enumerations, inline code/fenced code blocks for code/commands/identifiers
- Preserve every fact, name, number, quotation exactly as given - never invent, add, remove, or reorder
- Only restructure/lightly polish wording for flow/grammar
- Do not add introduction, conclusion, opinion, or commentary
- Do not repeat article title as heading, do not wrap whole answer in code fence
- Output only the reformatted article body as Markdown
-
Analyze the Input Text:
The input is a long article with a TL;DR, then body text. It contains technical details, quotes, numbers, model names, GitHub issues, etc. I need to parse it and reorganize into Markdown with headings and lists where appropriate, while preserving all content exactly.
Let me first carefully read and extract the structure. The text has natural sections/paragraphs. I need to identify topic shifts to place ## headings.
Let me copy the text and analyze:
Title: Ollama skips your JSON schema when a thinking model answers without thinking
Article:
TL;DR : Since 0.34.4, Ollama applies a format schema on a thinking model as a single grammar: anything until the thinking block closes, then the schema. That grammar also allows the response to end before the closing, and it treats everything up to the closing as thinking. So when a model like Gemma 4 decides not to think and answers straight away, its answer is "text before the closing" and nothing constrains it. On gemma4:e2b with think: true and a one-field schema, every answer that skipped thinking came back as bare text ( 391 , Paris ) with HTTP 200 and nothing in the log. Every answer that thought first matched the schema. With think: false the reply had the right shape every time. The open fix, ollama/ollama#18783 , made all 16 replayed requests match the schema, though not by constraining the direct answer: the model thought first every time instead. Until a release includes it, validate format replies in the client. The toolkit's check-ollama-format-think.sh tells you whether a model on your server does this. Upstream report: ollama/ollama#18774 . The symptom Debian 13 LXC, CPU only, Ollama 0.35.1 (the current release) from the release tarball, gemma4:e2b from the library. The report used gemma4:e4b ; the smaller model is what a CPU-only LXC runs at a usable speed. The schema asks for one string field: { "type" : "object" , "properties" : { "answer" : { "type" : "string" }}, "required" : [ "answer" ]} The report's request on /api/chat at temperature 0, with seeds 1 and 2, then with think: false : { "think" : true , "eval_count" : 4 , "thinking_chars" : 0 , "content" : "391" } { "think" : true , "eval_count" : 248 , "thinking_chars" : 607 , "content" : "{ \n \" answer \" : \" >391 \"\n }" } { "think" : false , "eval_count" : 11 , "thinking_chars" : 0 , "content" : "{ \" answer \" : \" ) \"\n }" } The first line is the bug. 391 parses as JSON, but it is a number, not an object with an answer field. The request returned HTTP 200 with done_reason: "stop" after four tokens, so it didn't run out of tokens. The second line is the same request with a different seed: the model thought first, and the reply has the right shape. /api/generate gave the same results. To see how often it happens, three prompts went out eight times each at the default temperature, seeds 0 to 7, on /api/chat : "What is 17 * 23? Answer with the number." thought 5, valid 5direct 3, valid 0 ("391") "What is the capital of France? One word." thought 3, valid 3direct 5, valid 0 ("Paris") "Is 91 prime? Answer yes or no." thought 8, valid 8direct 0 How often the model skips thinking depends on the question, but the result once it does is the same: 16 of 16 replies that thought first matched the schema, and 8 of 8 that skipped thinking didn't. In the report's think: false control, three requests on each endpoint at temperature 0, every reply had the right shape. Why it's easy to miss Whether you see it depends on the question. A test prompt that makes the model reason, which is what most people reach for when trying out a thinking model, never shows it. The prime question above went through thinking all eight times. The questions that trigger it are short and easy, which is also what most structured-output calls look like: classify this, extract that, yes or no. Nothing reports it either. The status is 200, done_reason is stop , and the server log has no line about the format. A reply that is only a word or a number can even look like a parsing problem on the client side. think: false working makes it look like think: true is just unsupported with format . It isn't: with thinking enabled the schema is enforced, but only on the part of the reply after the thinking. What's really going on #18479 ("apply structured outputs in a single pass on thinking models", merged 2026-09-22, first released in 0.34.4) made the format and the thinking one grammar. In llm/llama_server.go in 0.35.1, when the model's parser reports a thinking-close string and the request has a format, the schema is converted to a grammar and wrapped: if len ( req . ThinkingClose ) > 0 && ( lsReq . Grammar != "" || lsReq . JsonSchema != nil ) { ... lsReq . Grammar = thinkingGrammar ( req . ThinkingClose , lsReq . Grammar ) } The docstring of thinkingGrammar in llm/gbnf.go says what that wrapper allows: // thinkingGrammar returns a grammar that leaves the text before any of the // closings unconstrained, then constrains what follows the first complete // closing to the root of the format grammar. The response may end before any // closing. Gemma 4 opens thinking with <|channel> and closes it with <channel|> , and its parser reports <channel|> as the closing whenever thinking is enabled. The grammar assumes the response starts inside the thinking block. When the model starts with <|channel> , that holds: everything up to <channel|> is free, and the schema applies after it. When the model skips thinking and writes 391 followed by end of sequence, the grammar sees text before any closing, followed by an end the grammar allows. Nothing in it says the answer itself has to match the schema. The #18479 description acknowledges the limitation: "A model that answers without thinking stays unconstrained." With think: false the parser reports no closing, so the schema grammar is applied on its own from the first token, which is why that control always has the right shape. What the open fix does #18783 adds the opening string. For a model whose response can begin in content, the grammar becomes "either the schema, or <|channel> , then free thinking, then a required <channel|> , then the schema". The thinking branch can no longer end before the closing. The PR head ( 4121ed65 ) and its merge base ( 42e911bc ) were built in the same LXC, and each build got the eight prompt and seed pairs that had answered directly in the run above, on both endpoints: 0.35.1merge basePR #18783 answered directly15 of 1615 of 160 of 16 ... schema violated1515- thought first1 1 16 ... schema matched 1 1 16 eval_count, direct 2 to 42 to 4- eval_count, thought155 155 155 112 to 267 The one request that thought on 0.35.1 and the merge base (17 × 23, seed 1, /api/chat ) had answered directly in the earlier run; at the default temperature the same seed doesn't always take the same branch. The merge base behaved exactly like the release. On the PR build every request matched the schema. None of them was a constrained direct answer, though. The grammar no longer allows Paris as a first token, and of the two things it does allow, the start of the JSON object or <|channel> , the model picked <|channel> every time and thought for 260 to 634 characters. A capital-of-France call that took 2 tokens now takes about 113. That is a fair trade for getting the shape you asked for, but on a CPU box it is roughly fifty times the tokens for the questions that used to be the cheapest. The fix also does nothing for what goes inside the field. Across the builds, the 17 × 23 answers that went through thinking came back as "$391" , ">391" , "))391" and ")), " . That is this small model under a grammar, not the bug, but it means a reply that matches the schema still needs checking. The fix Until a release includes #18783: Validate the reply against the schema in the client. A reply that fails validation and has an empty thinking field is this case. On 0.35.1 every reply that thought first had the right shape, so a retry that lands in thinking will too; on prompts the model always answers directly, a retry won't help. Send think: false for calls where you need the shape and don't need the thinking. It had the right shape every time here. On gemma4:e2b the content was poor ( ")" , and in the report "\\text{391}" ), so check whether your model's answers survive it. Either way, don't assume that format guarantees the shape just because the request succeeded. With thinking enabled, it only covers the part of the reply after the thinking. What's really going on #18479 ("apply structured outputs in a single pass on thinking models", merged 2026-09-22, first released in 0.34.4) made the format and the thinking one grammar. In llm/llama_server.go in 0.35.1, when the model's parser reports a thinking-close string and the request has a format, the schema is converted to a grammar and wrapped: if len(req.ThinkingClose) > 0 && (lsReq.Grammar != "" || lsReq.JsonSchema != nil) { ... lsReq.Grammar = thinkingGrammar(req.ThinkingClose, lsReq.Grammar) } The docstring of thinkingGrammar in llm/gbnf.go says what that wrapper allows: // thinkingGrammar returns a grammar that leaves the text before any of the // closings unconstrained, then constrains what follows the first complete // closing to the root of the format grammar. The response may end before any // closing. Gemma 4 opens thinking with <|channel> and closes it with <channel|> , and its parser reports <channel|> as the closing whenever thinking is enabled. The grammar assumes the response starts inside the thinking block. When the model starts with <|channel> , that holds: everything up to <channel|> is free, and the schema applies after it. When the model skips thinking and writes 391 followed by end of sequence, the grammar sees text before any closing, followed by an end the grammar allows. Nothing in it says the answer itself has to match the schema. The #18479 description acknowledges the limitation: "A model that answers without thinking stays unconstrained." With think: false the parser reports no closing, so the schema grammar is applied on its own from the first token, which is why that control always has the right shape. What the open fix does #18783 adds the opening string. For a model whose response can begin in content, the grammar becomes "either the schema, or <|channel> , then free thinking, then a required <channel|> , then the schema". The thinking branch can no longer end before the closing. The PR head ( 4121ed65 ) and its merge base ( 42e911bc ) were built in the same LXC, and each build got the eight prompt and seed pairs that had answered directly in the run above, on both endpoints: 0.35.1merge basePR #18783 answered directly15 of 1615 of 160 of 16 ... schema violated1515- thought first1 1 16 ... schema matched 1 1 16 eval_count, direct 2 to 42 to 4- eval_count, thought155 155 155 112 to 267 The one request that thought on 0.35.1 and the merge base (17 × 23, seed 1, /api/chat ) had answered directly in the earlier run; at the default temperature the same seed doesn't always take the same branch. The merge base behaved exactly like the release. On the PR build every request matched the schema. None of them was a constrained direct answer, though. The grammar no longer allows Paris as a first token, and of the two things it does allow, the start of the JSON object or <|channel> , the model picked <|channel> every time and thought for 260 to 634 characters. A capital-of-France call that took 2 tokens now takes about 113. That is a fair trade for getting the shape you asked for, but on a CPU box it is roughly fifty times the tokens for the questions that used to be the cheapest. The fix also does nothing for what goes inside the field. Across the builds, the 17 × 23 answers that went through thinking came back as "$391" , ">391" , "))391" and ")), " . That is this small model under a grammar, not the bug, but it means a reply that matches the schema still needs checking. The fix Until a release includes #18783: Validate the reply against the schema in the client. A reply that fails validation and has an empty thinking field is this case. On 0.35.1 every reply that thought first had the right shape, so a retry that lands in thinking will too; on prompts the model always answers directly, a retry won't help. Send think: false for calls where you need the shape and don't need the thinking. It had the right shape every time here. On gemma4:e2b the content was poor ( ")" , and in the report "\\text{391}" ), so check whether your model's answers survive it. Either way, don't assume that format guarantees the shape just because the request succeeded. With thinking enabled, it only covers the part of the reply after the thinking. The generalisable habit When a constraint is applied to "the part after X", find out what happens when X never appears. Here the constraint was written for responses that begin with thinking, and a response with no thinking at all fell through it. It is the second structured-output gap in Ollama 0.35 here in a week. On one of its two chat paths, your schema's keys are re-sorted alphabetically before the grammar sees them. Both times the reply was HTTP 200 and nothing in the log said anything was wrong. The toolkit's check-ollama-format-think.sh sends a few short prompts to a model with think: true and a schema, and reports how many replies skipped thinking and how many of those broke the schema. Toolkit This post's fix is available as a tested, ready-to-run script in the toolkit. See the toolkit →
Wait, the input has "Title:" and "Article:" prefixes, but the rule says "Do not repeat the article title as a heading and do not wrap the whole answer in a code fence." So I should output just the article body reformatted, not including the "Title:" line as a heading. The title is given but I shouldn't output it as a heading. I'll just output the reformatted article body.
Also, the text has some duplication near the end - it seems like the latter part is repeated. I need to preserve every fact, name, number, quotation exactly as given. So I must include all text, but I can structure it. The duplication might be part of the original; I should keep it as-is, but maybe I can avoid duplicating huge chunks if it's truly a copy-paste error in the prompt. However, the rule says "Preserve every fact, name, number, and quotation exactly as given. Never invent, add, remove, or reorder information." So I must include all content. I'll treat the entire text after "Article:" as the body, and I need to reformat it.
Let me re-read the prompt: "You are given the raw body of a news article that arrived as an undifferentiated wall of text. Reformat it into clean, readable Markdown for a technical audience." Then rules. The input includes "Title: ..." and "Article:". I should probably strip the "Title:" and "Article:" labels, as they're just metadata for me, not part of the article body to reformat. But the rule says "Do not repeat the article title as a heading". So I must not output a heading with the title. I'll just output the article body.
Let me examine the text structure. It has a TL;DR, then several paragraphs. I need to break into short paragraphs, add headings where topic shifts. I'll go through and identify natural sections.
I'll outline the content mentally:
- TL;DR summary paragraph
- Description of the bug with examples (schema, thinking, gemma4)
- Report details: symptom, Debian, Ollama version, model, schema
- Specific request examples (the three lines with think true/false)
- Frequency analysis: three prompts x 8 seeds
- Why it's easy to miss
- What's really going on (reference to #18479)
- ThinkingGrammar docstring explanation
- How Gemma 4 thinking blocks work
- The open fix #18783
- PR head/merge base results
- The fix recommendation
- Generalisable habit
- Second structured-output gap mention
- Toolkit script mention
Some sections are repeated at the end, which might be a artifact. I'll include everything exactly as given, but I'll try to avoid duplicating large verbatim blocks if it's clearly a repeat. However, to be safe, I'll include all text, but structure it so the repetition is evident but not wasted space. Actually, the rule says "Never ... reorder information." So I must keep the order. I'll just format what's there.
Let me carefully copy the article text without the "Title:" and "Article:" prefixes. The input starts with "Title: ..." then "Article:". I'll assume the article body starts after "Article:". I'll strip those lines.
Now, I need to produce Markdown. I'll create headings like ## TL;DR, ## The Bug, ## Report Details, ## Why It's Easy to Miss, ## What's Really Going On, ## The Open Fix #18783, ## The Fix, ## Generalisable Habit, etc.
I need to ensure short paragraphs separated by blank lines. I'll break long paragraphs. I'll use bullet lists for enumerations where appropriate, but must preserve order and content.
Let me start processing text segment by segment.
First segment: "TL;DR : Since 0.34.4, Ollama applies a format schema on a thinking model as a single grammar: anything until the thinking block closes, then the schema. That grammar also allows the response to end before the closing, and it treats everything up to the closing as thinking. So when a model like Gemma 4 decides not to think and answers straight away, its answer is "text before the closing" and nothing constrains it. On gemma4:e2b with think: true and a one-field schema, every answer that skipped thinking came back as bare text ( 391 , Paris ) with HTTP 200 and nothing in the log. Every answer that thought first matched the schema. With think: false the reply had the right shape every time. The open fix, ollama/ollama#18783 , made all 16 replayed requests match the schema, though not by constraining the direct answer: the model thought first every time instead. Until a release includes it, validate format replies in the client. The toolkit's check-ollama-format-think.sh tells you whether a model on your server does this. Upstream report: ollama/ollama#18774 ."
I'll make a ## TL;DR heading, and put that paragraph.
Next: "The symptom Debian 13 LXC, CPU only, Ollama 0.35.1 (the current release) from the release tarball, gemma4:e2b from the library. The report used gemma4:e4b ; the smaller model is what a CPU-only LXC runs at a usable speed. The schema asks for one string field: { "type" : "object" , "properties" : { "answer" : { "type" : "string" }}, "required" : [ "answer" ]}"
This is a section. I'll maybe heading ## The Symptom.
Then: "The report's request on /api/chat at temperature 0, with seeds 1 and 2, then with think: false : { "think" : true , "eval_count" : 4 , "thinking_chars" : 0 , "content" : "391" } { "think" : true , "eval_count" : 248 , "thinking_chars" : 607 , "content" : "{ \n " answer " : " >391 "\n }" } { "think" : false , "eval_count" : 11 , "thinking_chars" : 0 , "content" : "{ " answer " : " ) "\n }" } The first line is the bug. 391 parses as JSON, but it is a number, not an object with an answer field. The request returned HTTP 200 with done_reason: "stop" after four tokens, so it didn't run out of tokens. The second line is the same request with a different seed: the model thought first, and the reply has the right shape. /api/generate gave the same results."
This has examples; I can format the JSON as code blocks inline or fenced. The rule says "Use inline code or fenced code blocks where code, commands, or identifiers appear." So I can use fenced code blocks for those JSON lines. But I need to preserve exact content. I'll use fenced code blocks with the JSON, but ensure they're rendered as code. However, the rule also says "Preserve every fact, name, number, and quotation exactly as given." Using a fenced code block preserves it exactly. I'll do that.
But careful: the text has those JSON objects on separate lines. I can put them in a fenced block. However, the text also has surrounding prose. I'll integrate.
Next: "To see how often it happens, three prompts went out eight times each at the default temperature, seeds 0 to 7, on /api/chat : "What is 17 * 23? Answer with the number." thought 5, valid 5direct 3, valid 0 ("391") "What is the capital of France? One word." thought 3, valid 3direct 5, valid 0 ("Paris") "Is 91 prime? Answer yes or no." thought 8, valid 8direct 0 How often the model skips thinking depends on the question, but the result once it does is the same: 16 of 16 replies that thought first matched the schema, and 8 of 8 that skipped thinking didn't. In the report's think: false control, three requests on each endpoint at temperature 0, every reply had the right shape."
I'll make a heading ## Frequency of Skipping Thinking.
Next: "Why it's easy to miss Whether you see it depends on the question. A test prompt that makes the model reason, which is what most people reach for when trying out a thinking model, never shows it. The prime question above went through thinking all eight times. The questions that trigger it are short and easy, which is also what most structured-output calls look like: classify this, extract that, yes or no. Nothing reports it either. The status is 200, done_reason is stop , and the server log has no line about the format. A reply that is only a word or a number can even look like a parsing problem on the client side. think: false working makes it look like think: true is just unsupported with format . It isn't: with thinking enabled the schema is enforced, but only on the part of the reply after the thinking."
Heading ## Why It's Easy to Miss.
Next: "What's really going on #18479 ("apply structured outputs in a single pass on thinking models", merged 2026-09-22, first released in 0.34.4) made the format and the thinking one grammar. In llm/llama_server.go in 0.35.1, when the model's parser reports a thinking-close string and the request has a format, the schema is converted to a grammar and wrapped: if len ( req . ThinkingClose ) > 0 && ( lsReq . Grammar != "" || lsReq . JsonSchema != nil ) { ... lsReq . Grammar = thinkingGrammar ( req . ThinkingClose , lsReq . Grammar ) } The docstring of thinkingGrammar in llm/gbnf.go says what that wrapper allows: // thinkingGrammar returns a grammar that leaves the text before any of the // closings unconstrained, then constrains what follows the first complete // closing to the root of the format grammar. The response may end before any // closing. Gemma 4 opens thinking with <|channel> and closes it with <channel|> , and its parser reports <
Comments
No comments yet. Start the discussion.