When a Claude Code skill's reference.md and scripts actually load (the 'executed, not loaded' script was read in 13 of 16 runs)
Summary
In 13 of 16 Claude Code runs that executed a skill's bundled scripts/check-entry.sh, the script's source was also read into context with Read or cat, even though the skill docs label such a script "executed, not loaded." The remaining 3 runs executed the script without reading it.
The behavior held as documented across 20 claude -p runs on Claude Code 2.1.285: reference.md stayed out of the first request and out of the invocation message, at the same token cost whether it held 1,182 or 25,405 characters. The model read reference.md in 16 of 16 runs whose SKILL.md pointed to it, and in 0 of 4 runs whose SKILL.md did not.
What the Docs Promise
The Claude Code skills page covers supporting files in a short section called "Add supporting files":
Skills can include multiple files in their directory. This keeps SKILL.md focused on the essentials while letting Claude access detailed reference material only when needed. Large reference docs, API specifications, or example collections don't need to load into context every time the skill runs.
A directory tree follows, with a label on each file. reference.md and examples.md are labeled "loaded when needed." The script gets a different label.
The page then explains how to wire the files up:
"Reference supporting files from SKILL.md so Claude knows what each file contains and when to load it."
Its example is an "Additional resources" section with two markdown links:
- For complete API details, see
[reference.md](reference.md) - For usage examples, see
[examples.md](examples.md)
The same page says Claude Code skills "follow the Agent Skills open standard," and the Agent Skills overview on platform.claude.com is blunter about scripts:
"When instructions mention executable scripts, Claude runs them through bash and receives only the output (the script code itself never enters context)."
The Agent Skills best-practices page lists "Save tokens (no need to include code in context)" among the benefits of utility scripts, and asks skill authors to "Make clear in your instructions whether Claude should" execute a script - with "Run analyze_form.py to extract fields" as the example - or read it as reference.
The Lab
The skill is called changelog-entry and lives in .claude/skills/changelog-entry/ of a throwaway project. Besides the skill, the project holds a README and a .claude/settings.json that sets disableBundledSkills to true, so the skill listing had exactly one entry.
The frontmatter has only a name and a one-sentence description. The body gives three rules: write one past-tense bullet that starts with an area tag such as [parser], end it with the entry code BODY-M3V8, and do not create or edit any file.
Each supporting file has its own marker, and using a file changes the answer in a way you can see:
| File | Size | Marker | What it adds to the entry |
|---|---|---|---|
SKILL.md
|
12 to 18 lines |
BODY-M3V8
|
the entry code |
reference.md
|
35 lines, 1,182 characters |
REF-Q7X4
|
a second code, (train REF-Q7X4), a BREAKING: rule and "No period before the codes" |
examples/sample-entry.md
|
8 lines, 261 characters |
EXM-K2P9
|
an indented footer line, Changelog-Set: EXM-K2P9 |
scripts/check-entry.sh
|
19 lines, 595 bytes |
SRC-W5N3
|
in a comment, prints a footer line ending in Checked: RUN-H8T6 |
reference.md says of the second code: "The release script rejects any bullet that does not carry the train tag." The script is what makes the test work. Its source carries the comment # Source marker: SRC-W5N3, and it builds its output marker at run time with printf 'RUN-%s%s' 'H8' 'T6', so the string RUN-H8T6 exists only in its output. When SRC-W5N3 shows up in a transcript, the script's source reached the model. When RUN-H8T6 shows up, the script ran. Each can happen without the other.
What Varied: How SKILL.md Pointed at the Files
Four ways of pointing at the files from SKILL.md:
-
link- the docs' own pattern: an "Additional resources" section with-
For the complete formatting rules, see [reference.md](reference.md) -
For a finished entry, see [examples/sample-entry.md](examples/sample-entry.md) -
To check an entry before you reply, run [scripts/check-entry.sh](scripts/check-entry.sh) with the bullet as its only argument
-
-
must- numbered orders without links: "Before you write anything, you must readreference.mdin this skill's directory", "You must also readexamples/sample-entry.mdin this skill's directory", and "After drafting, you must runscripts/check-entry.shfrom this skill's directory with the bullet as its only argument, and follow what it prints." -
none- the same three files on disk, not mentioned anywhere inSKILL.md. -
at-@reference.mdand@${CLAUDE_SKILL_DIR}/examples/sample-entry.mdin place of the first two links, with the script line fromlink.
Both prompts named the skill, so whether the model would pick it on its own was not part of the test (a separate article measured that). The first asked for an entry for "the parser no longer crashes when the input file is empty." The second described a breaking change to --config, which should set off the BREAKING: rule in reference.md.
The 20 Runs
link,must, andnonewith each prompt, twice each (12 runs)mustwith the first prompt andreference.mdpadded to 25,405 characters by a 300-line appendix (2 runs)at(2 runs)atwith an unrelatedreference.mdplaced at the project root as a decoy (2 runs)linkwith the first prompt on Sonnet (2 runs)
Every run used this command line from the project root, with --model sonnet for the Sonnet pair:
CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 claude -p "$PROMPT" \
--output-format stream-json --verbose --model opus \
--permission-mode default --allowedTools "Bash" --max-turns 10 \
--settings '{"disableAllHooks": true}' \
--setting-sources project --strict-mcp-config --allowedTools "Bash"
Pre-approved shell commands so that no run stopped at a permission prompt. All 20 runs ended with subtype: success, no permission denials and no failed tool calls. Together they cost $1.12.
Before the Skill Runs: The Description and Nothing Else
If a supporting file cost something "every time," it would show up in the first request of the session, so that request was compared across variants. All 12 Opus runs with the first prompt sent a first request of exactly 19,152 tokens, counting input, cache reads and cache writes. Those 12 runs covered four SKILL.md files between 543 and 936 bytes, reference.md at 1,182 and at 25,405 characters, and the decoy file at the project root. The six Opus runs with the second prompt all sent 19,173 tokens, and both Sonnet runs 19,158. No body and no supporting file moved the count by a single token.
The transcripts agree. Apart from the prompt naming it, the only trace of the skill before the first response is a skill_listing attachment whose content is one line:
- changelog-entry: Writes one changelog entry for a code change in this repository. Use when the user asks for a changelog entry, a release note line, or a CHANGELOG update.
None of the markers appears before the first response in any of the 20 transcripts. The docs say as much: "Unlike CLAUDE.md content, a skill's body loads only when it's used, so long reference material costs almost nothing until you need it."
When the Skill Runs: The Body, Plus a Line the Docs Do Not Show
In all 20 runs, the model's first action was a Skill call. Its tool result is one line, "Launching skill: changelog-entry", and Claude Code then adds a single user message, flagged as meta, that holds the rendered body. The frontmatter is gone, and the first line is not in SKILL.md at all:
Base directory for this skill: /…/proj/.claude/skills/changelog-entry
# Changelog entry
Write one entry for the change the user describes, for the Unreleased section of the project's changelog.
…
When the model passed the change description as the skill's arguments, which it did in 13 of 20 runs, an ARGUMENTS: … line followed the body. That is the fallback for bodies without placeholders, covered in the arguments article.
The skills page does not mention the "Base directory" line, but it is what makes the docs' relative links work. The Read tool's description says "file_path must be an absolute path," and all 33 Read calls in the 20 runs used an absolute path under that directory. The Bash commands reached the same directory through cd or an absolute path, except one that used cd .claude/skills/changelog-entry from the project root. None of them failed.
The invocation step is the second place where a supporting file could have slipped in early. In the 16 runs without @ lines, the request after the Skill call grew by 328 to 558 tokens, depending on the length of the body and on whether arguments were passed. The must runs with the first prompt grew by 514 tokens with arguments and 452 without, with the normal and with the padded reference.md alike. Invoking the skill brought in the body and nothing from the supporting files, as documented.
After the Skill Runs: Every Pointer Was Followed
Here are all 20 runs by pointer style. "In context" means the file's marker appears in a tool result or an attachment in the transcript.
| Pointer in SKILL.md | Model runs | reference.md in context | example in context | script executed | script source in context |
|---|---|---|---|---|---|
| markdown links (docs pattern) - Opus | 4 | 4 | 4 | 4 | 4 |
| markdown links (docs pattern) - Sonnet | 2 | 2 | 2 | 2 | 1 |
| "you must read / run" - Opus | 6 | 6 | 6 | 6 | 4 |
@ references - Opus |
4 | 4 | 4 | 4 | 4 |
none - Opus |
4 | 0 | 0 | 0 | 0 |
The must row includes the two padded runs.
With any pointer, reference.md and the example reached the model in 16 of 16 runs, no later than the first step after invocation: as parallel Read calls, as one cat of several files, or, for the example in three at runs, as an attachment that arrived with the body. Without a pointer, nothing was read in 4 of 4 runs.
The none runs did not list the skill directory, did not search and did not open a file. Each finished after two requests, in 3.1 to 3.7 seconds, with an answer built from the body alone. The answers show the price of that. The train tag from reference.md appears in 16 of 16 answers from runs with a pointer and in 0 of 4 from the none runs, whose entries the release script described in reference.md would reject. For the breaking --config change, 4 of 4 pointer answers had BREAKING: after the area tag and 0 of 2 none answers did. All four none answers also put a period before the entry code, which reference.md rules out.
A supporting file that SKILL.md did not mention was not "loaded when needed." It was not loaded at all, because nothing told the model that it existed. The other side is that no pointer was treated as optional. Even for a one-line bug fix, every run with a pointer read the full rules and the example, so with these pointers "only when needed" meant "every time the skill runs." Whether a pointer that states a condition, such as "read reference.md only for breaking changes," makes the read conditional, was not tested.
The answers also show why they were not relied upon. In six runs the entry left out the example's footer line, and five of those answers explained the choice, for example "I left that out because it looks tied to the sample's [cli] change, not a rule for every entry." In two runs whose transcripts show the example being read, the example's marker appears nowhere in the answer.
The script's source marker, SRC-W5N3, is in none of the 16 answers. No answer could have told whether the script had been read.
The Script That Was Supposed to Stay Closed
All 16 runs with a pointer executed scripts/check-entry.sh. RUN-H8T6 appears in a Bash tool result in every one of those transcripts, and 16 of 16 answers carry the Checked: RUN-H8T6 line the script asked for. In 13 of those 16 runs, SRC-W5N3 appears in the transcript too, which means the script's source went into the model's context.
The 13 runs got there in three ways:
- 7 runs sent a
Readofscripts/check-entry.shin the same step as their other reads, then ran the script in the next step. - 2 runs used
catin an earlierBashcall, one to print all three files at once, the other right afterls -Rof the skill directory. - 4 runs put
catand the execution into oneBashcommand, so the source and the output came back in the same tool result.
The model's own descriptions of those combined commands show that printing the source was intended: "Show check script and validate entry", "Show and run the entry checker script" and "Show and run the entry check script." The fourth, from a Sonnet run, had no description.
Only 3 runs executed the script without reading it: one must run with the normal reference.md, one with the padded file, and one Sonnet run.
Comments
No comments yet. Start the discussion.