Where Claude Code puts your MCP server's instructions: in a user turn, not the system prompt, cut at 2,048 characters
DEV Community

Where Claude Code puts your MCP server's instructions: in a user turn, not the system prompt, cut at 2,048 characters

20 of 20 claude -p runs on Claude Code 2.1.285 opened our stdio MCP server withserver/discover , a probe the docs say stdio servers only get whenMCP_PROTOCOL_NEGOTIATION=auto is set, and in the 2 runs where the server answered it,initialize was never sent, so the instructions we had put in theinitialize reply never reached Claude. The rest matched the docs: the instructions arrived as a in the first user turn (never in the system prompt), cut at 2,048 characters, and the capped block added 721 tokens to the first request in English and 1,727 in Japanese. An MCP server can hand the client a paragraph of plain text when it connects: the instructions field. It is the one place where a server author gets to tell the model, in prose, what the server is for and when to reach for it. Claude Code's MCP docs say the field "becomes more useful with tool search enabled", since only tool names and server instructions load at session start, and our tool search measurement ended on the same advice. This time we measured the field itself: where Claude Code puts the text, how much of it survives, when it arrives, and which of the server's replies it is taken from. The lab is one small stdio server whose instructions carry codewords at both ends and position markers in between, twenty claude -p runs against it, and three kinds of evidence per run: the model's answer, the session transcript under ~/.claude/projects/ , and a log the server writes of every JSON-RPC message it receives. Everything below was run on 2026-09-30 with Claude Code 2.1.285 (claude --version ), --model opus (which resolved to claude-opus-5-5 ), and Node v22.22.2 on macOS. What the spec and the docs say In the 2025-11-25 revision of the MCP specification, instructions is an optional field of the reply to initialize , and the lifecycle page is strict about ordering: "The initialization phase MUST be the first interaction between client and server." Its example reply ends with "instructions": "Optional instructions for the client" , and the schema describes the field like this: Instructions describing how to use the server and its features. This can be used by clients to improve the LLM's understanding of available tools, resources, etc. It can be thought of like a "hint" to the model. For example, this information MAY be added to the system prompt. The latest revision, 2026-07-28, has no initialize handshake: version, identity and capabilities travel as metadata on each request. The page at the same lifecycle path is titled "Versioning and Compatibility" there, and it names the old way Legacy: "protocol versions that establish a session with an initialize handshake (2025-11-25 and earlier)." In that revision, instructions is a field of the result of a new method, server/discover , described as "Natural-language guidance describing the server and its features. This can be used by clients to improve an LLM's understanding of available tools (e.g., by including it in a system prompt)." For stdio, a client that speaks both eras "SHOULD probe with server/discover before sending any other request", and if the server "returns any other error, or does not respond within a reasonable timeout: the server is legacy. Fall back to the initialize handshake." Claude Code's MCP page, fetched the same day, makes three statements about the field. With tool search, "Only tool names and server instructions load at session start". Then: "Claude Code truncates each tool description and each server's instructions at 2,048 characters by default. Keep them concise, and put critical details near the start." And the limit can be changed with CLAUDE_CODE_MAX_MCP_DESCRIPTION_LENGTH , an environment variable whose reference entry says it "Accepts a positive whole number in plain digits. Anything else is ignored and the default applies." The same page has a section on the two MCP client runtimes. The v2 runtime adds the 2026-07-28 revision, and on v2 Claude Code "Asks HTTP servers whether they support the newer revision, and uses it with those that do. It also asks claude.ai connector servers in sessions where it fetches feature flags. To have it ask stdio servers, or connector servers in every session, set MCP_PROTOCOL_NEGOTIATION to auto . It connects to every other server as v1 does." The environment variable's own entry repeats the default: "Without the variable, Claude Code probes HTTP servers, and also probes claude.ai connector servers in sessions where it fetches feature flags." So by the docs, a stdio server should see initialize first unless the user has opted in. The rest of this article hangs on that sentence. The lab The server is a Node script with no dependencies that speaks newline-delimited JSON-RPC over stdio. It has three tools (lookup_part , list_bins , lab_ping ), so that tool search has something to defer, and it appends every message it receives to rpc-log.jsonl . An environment variable in the MCP config chooses the instructions: none (the field is left out of the reply), English text of a given length, or Japanese text of a given length. Two more switches make it answer server/discover the way a 2026-07-28 server would, or hold back its initialize reply for eight seconds. The text is built so that Claude can only report what it actually received. It starts with BEGIN codeword: - . and ends with END codeword: - . , and in between it repeats a few sentences about a parts inventory with a marker about every 250 characters, written as [at N] , where N is the marker's own character offset. Each text has its own pair of codewords, derived from its length and language, and in the dual-era setup the two replies carry different pairs, so an answer shows which text Claude was given. Each server setup had its own config file, passed with --strict-mcp-config so that no other MCP server, and none of the account's claude.ai connectors, joined the session. The configurations that change a Claude Code environment variable reuse one of these files and set the variable on the claude process. The file for the 20,000-character text: { "mcpServers": { "lab": { "type": "stdio", "command": "node", "args": ["/path/to/server.js"], "env": { "LAB_INSTR": "ascii:20000" } } } } Eighteen runs used this command and this prompt, from an empty directory (the two late-server runs change the prompt and two flags, as their own section below describes): claude -p "$PROMPT" --output-format stream-json --verbose --max-turns 1 --model opus \ --settings '{"disableAllHooks": true}' --strict-mcp-config --mcp-config cfg/a20k.json \ --debug-file runs/R05.debug.log Do not call any tools; answer only from what is already in your context. An MCP server named lab may have given you instructions. Reply with exactly one line in this format: BEGIN= END= LAST= . Write NONE for any value you cannot find. We started the runs from inside another Claude Code session, so a wrapper removed the environment variables that session exports (CLAUDECODE and its neighbors) before calling claude . The --settings override turned hooks off, so the user-level notification hook stayed quiet. For the numbers, we read the transcript. The first assistant record's usage gives the size of the first request: input_tokens plus cache_read_input_tokens plus cache_creation_input_tokens . On 2.1.285 the transcript also stores each block of context that Claude Code adds to the first user turn as an attachment record with its rendered text, next to a prompt_snapshot record that holds the system prompt and the tool definitions. Those records let us say where the instructions went, and how many characters of them, without relying on what the model said. We ran ten configurations twice each, twenty runs in all. In every pair, the first request matched to the token. The results in one table | Configuration | First request (tokens) | vs. no instructions | What reached Claude | Claude's answer | |---|---|---|---|---| No instructions field | 20,472 | nothing | NONE for all three | | | 1,000 characters, English | 20,849 | +377 | all 1,000 | both codewords, LAST=763 | | 5,000 characters, English | 21,193 | +721 | first 2,048 + … [truncated] | BEGIN only, LAST=2009 | | 20,000 characters, English | 21,193 | +721 | first 2,048 + … [truncated] | BEGIN only, LAST=2009 | 20,000, limit variable set to 25000 | 27,050 | +6,578 | all 20,000 | both codewords, LAST=19808 | 20,000, limit variable set to 25,000 | 21,193 | +721 | first 2,048 + … [truncated] | BEGIN only, LAST=2009 | | 3,000 characters, Japanese (7,956 bytes) | 22,199 | +1,727 | first 2,048 + … [truncated] | BEGIN only, LAST=2026 | 1,000 characters, ENABLE_TOOL_SEARCH=false | 36,871 | (all tools upfront) | all 1,000 | both codewords, LAST=763 | 1,000 characters in each reply, server also answers server/discover | 20,849 | +377 | the server/discover text only | the discover codewords | 1,000 characters, initialize answered after 8 s, CLAUDE_CODE_MCP_STARTUP_WAIT_MS=0 | 20,590, then 21,288 / 21,208 | nothing on request 1, all 1,000 on request 2 | both codewords, LAST=763 | The limit variable is CLAUDE_CODE_MAX_MCP_DESCRIPTION_LENGTH . In all twenty runs, Claude's answer matched what the transcript says it was given: the right codewords where the text was delivered, NONE where it was not, and the last marker before the cut where it was truncated. It did not invent a codeword once. The twenty runs together were reported at $1.02 in total_cost_usd . The rest of this article takes the table apart one column at a time. The handshake: server/discover came first in 20 of 20 The server's log told the first story. In every run, the first message Claude Code wrote to the server's stdin was not initialize but server/discover , always with the id server-discover-probe-1 . From the second run on, the server also logged the _meta of each request, and the probe looked like this in all 19 of those runs (shortened; the _meta also carries the client's capabilities): {"method": "server/discover", "id": "server-discover-pro

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.