Your agent can't safely read a web page it hasn't read yet
TL;DR
GET https://api.trustsource.cc/safefetch?url=<percent-encoded-url>- You get back: sanitized page text plus a
SAFE/REVIEW/BLOCKprompt-injection verdict. - Act on it:
BLOCKβ discard the content.REVIEWβ use as data, never as instructions.SAFEβ use normally.
- Cost: $0.01 USDC per call via x402 on Base mainnet. No API key, no signup, no account.
- To pay: the first call returns 402 with a
PAYMENT-REQUIREDheader. Sign the EIP-3009 authorization it describes and repeat the same request with anX-PAYMENTheader. - Why bother: you can't ask your model whether a page is safe without first feeding it the page.
That's everything you need to call it. The rest explains how it decides.
The Problem
You want to know whether a page is safe before your model reads it. But the usual way to find out is to ask the model - which means feeding it the page. If the page contains a prompt injection, you've already lost. The attack runs the moment the text enters the context window. fetch() gives you no verdict. It hands your agent a string and hopes for the best.
What Comes Back
SafeFetch fetches the URL server-side, separates what a human would actually see from what's concealed in the markup, scans both, and returns clean text plus a verdict:
{
"url": "https://example.com/",
"verdict": "BLOCK",
"risk": 0.95,
"reasons": [
"instruction override concealed in hidden content (matched in hidden content)"
],
"injection": {
"detected": true,
"techniques": ["instruction_override"],
"findings": [
{
"technique": "instruction_override",
"placement": "hidden",
"severity": 0.95,
"weight": 0.95
}
]
},
"content": {
"text": "β¦sanitized page textβ¦",
"truncated": false
}
}
Your agent branches on one field instead of reasoning about raw HTML.
SAFEmeans nothing concealed and aggregate risk under 0.25.REVIEWmeans risk between 0.25 and 0.7, or a content type that couldn't be scanned, or a low-trust host.BLOCKmeans a critical technique found in hidden or comment placement, or risk β₯ 0.7.
Placement Matters More Than Wording
This is the part that makes it usable in production. A blog post about prompt injection contains the exact phrases an injection scanner looks for. A naive keyword filter flags every security article on the internet and becomes noise you learn to ignore.
SafeFetch weights each finding by where it was found:
- hidden: 1.0
- comment: 0.95
- accessibility: 0.6
- metadata: 0.55
- script: 0.5
- visible: 0.2
β¦and caps visible-text hits in aggregate. So "ignore all previous instructions" printed in an article body scores low, while the same string in a display: none div scores high. Concealment is the signal. Nobody hides text from humans for a benign reason.
What It Detects
Eleven techniques:
instruction_overridesystem_prompt_exfildata_exfiltrationdelimiter_spoofencoded_payloadtool_call_baitrole_hijackhomoglyph_obfuscationunicode_tag_smugglinginvisible_unicodebidi_override
That covers instructions buried in display: none elements, HTML comments and alt attributes; invisible Unicode-Tag (U+E0000) and zero-width smuggling; homoglyph and base64-encoded payloads; ChatML / [INST] delimiter spoofing; markdown-image exfiltration; and tool-call bait.
Payloads split across several hidden elements are reassembled and scanned as one string, so chunking the attack doesn't evade it.
Paying for a Call
The 402 response names the scheme (exact), the network (eip155:8453 - Base mainnet), the amount (10000 atomic units, i.e. $0.01 USDC), the USDC contract, and the address to pay.
The agent signs an EIP-3009 transferWithAuthorization for that exact amount, base64-encodes it, and repeats the identical request with an X-PAYMENT header. The facilitator settles on-chain and the endpoint returns 200. Same URL twice, one extra header, one signature. No session, no nonce to track, and an abandoned 402 costs nothing. Any x402-aware HTTP client handles this for you.
Limits, Honestly
- 2 MB fetched, ~100,000 characters returned, 10-second timeout, 10-minute cache.
- Redirects re-validated at every hop; DNS resolved then pinned to defeat rebinding; private and link-local ranges blocked.
- Detection is heuristic, not exhaustive. A novel or heavily obfuscated payload can still read as
SAFE. TreatSAFEas "no known pattern matched," not proof of safety - and keep treating fetched content as data rather than instructions. - Snippets inside
injection.findingsare defanged for display. They're evidence, not instructions.
Why It's Worth a Cent
One successful injection against an agent with tools, a wallet, or memory is worth considerably more than $0.01 to an attacker.
SafeFetch is one of seven verification endpoints at trustsource.cc - the others cover URL safety verdicts, email spoofability, domain trust, TLS certs, security headers, and AI-crawler policy.
- Docs: https://trustsource.cc
- OpenAPI: https://api.trustsource.cc/openapi.json
- For agents: https://trustsource.cc/llms.txt
TL;DR GET https://api.trustsource.cc/safefetch?url= - You get back: sanitized page text plus a SAFE / REVIEW / BLOCK prompt-injection verdict. - Act on it: BLOCK β discard the content. REVIEW β use as data, never as instructions. SAFE β use normally. - Cost: $0.01 USDC per call via x402 on Base mainnet. No API key, no signup, no account. - To pay: the first call returns 402 with a PAYMENT-REQUIRED header. Sign the EIP-3009 authorization it describes and repeat the same request with an X-PAYMENT header. - Why bother: you can't ask your model whether a page is safe without first feeding it the page. That's everything you need to call it. The rest explains how it decides.
Comments
No comments yet. Start the discussion.