I Tested Cloudflare Workers AI Free Tier and Found 24 Powerful LLMs You Can Use Today
Most developers know about OpenAI, Anthropic, Groq, and Cerebras. What many developers don't realize is that Cloudflare Workers AI provides access to a large collection of open-source models through a single API endpoint. The interesting part? Many of these models are usable on Cloudflare's free tier. I recently decided to test Cloudflare Workers AI to answer three questions: Which models are actually available on the free plan? Which models are best for coding? How do Cloudflare's limits compare to Groq, Cerebras, and other providers? The results were honestly surprising. The Experiment Cloudflare exposes a model catalog API: Invoke-RestMethod -Uri "https://api.cloudflare.com/client/v4/accounts/$AccountId/ai/models/search" -Headers $Headers My account returned more than 300 models. Initially, I assumed that every model listed would be available. I was wrong. Some models appeared in the catalog but were blocked when inference requests were executed. For example: zai-org/glm-5.3 zai-org/glm-5.3-flash zai-org/glm-5.2 moonshotai/kimi-k2.6 moonshotai/kimi-k2.7-code returned: { "code": 5035, "message": "Model is not available on the Workers Free plan" } This led me to create a discovery script that: Enumerated all available models Executed a test inference Recorded successful responses Recorded paid-plan restrictions Models That Worked on the Free Tier After testing, these models successfully accepted inference requests. OpenAI openai/gpt-oss-20b openai/gpt-oss-120b Qwen qwen/qwen2.5-coder-32b-instruct qwen/qwen3-30b-a3b-fp8 qwen/qwen3.8-27b qwen/qwq-32b DeepSeek deepseek-ai/deepseek-r1-distill-qwen-32b Meta Llama meta/llama-3.1-8b-instruct-fp8 meta/llama-3.2-1b-instruct meta/llama-3.2-3b-instruct meta/llama-3.3-70b-instruct-fp8-fast meta/llama-4-scout-17b-16e-instruct Google Gemma google/gemma-2b-it-lora google/gemma-7b-it-lora google/gemma-4-26b-a4b-it aisingapore/gemma-sea-lion-v4-27b-it Mistral mistral/mistral-7b-instruct-v0.2-lora mistralai/mistral-small-3.1-24b-instruct ZAI @cf/zai-org/glm-4.7-flash IBM ibm-granite/granite-4.0-h-micro NVIDIA nvidia/nemotron-3-120b-a12b Biggest Surprise I initially started testing because I wanted access to GLM 5.3 Flash. The result? GLM 4.7 Flash โ GLM 5.2 โ GLM 5.3 โ GLM 5.3 Flash โ GLM 4.7 Flash was available on the free tier while all newer GLM models required a paid plan. Best Models For Coding After testing many providers over the last year, this is how I would rank the Cloudflare free models for software development. ๐ฅ GPT-OSS-120B @cf/openai/gpt-oss-120b Best overall coding model. Excellent at: Terraform Kubernetes DevOps AWS Architecture Refactoring Multi-file repositories If I could pick only one model, it would be GPT-OSS-120B. ๐ฅ Qwen2.5-Coder-32B @cf/qwen/qwen2.5-coder-32b-instruct Purpose-built coding model. Excellent at: Generating code Reviewing pull requests Fixing bugs Understanding repositories Continue.dev Cline This model consistently performs above its size class. ๐ฅ DeepSeek-R1-Distill-Qwen-32B @cf/deepseek-ai/deepseek-r1-distill-qwen-32b Best reasoning model. Excellent at: Root cause analysis Complex debugging Architecture reviews Agent workflows This is the model I would use when a pipeline breaks at 2 AM and nobody knows why. Honorable Mentions QWQ-32B @cf/qwen/qwq-32b Very strong reasoning. Llama 4 Scout @cf/meta/llama-4-scout-17b-16e-instruct Excellent balance of speed and intelligence. Llama 3.3 70B @cf/meta/llama-3.3-70b-instruct-fp8-fast Great for reviews, documentation, and system design. Understanding Cloudflare Limits One thing I wanted to understand was whether Cloudflare behaves like Groq. The answer is no. Groq commonly applies limits per model. For example: Model A → X RPM Model B → Y RPM Cloudflare works differently. Cloudflare's free tier provides: 10,000 Neurons per day This is effectively a daily AI budget shared across your account. Cloudflare also documents: 300 requests per minute for text generation workloads. So the practical model looks like: Account Level Daily Budget + Text Generation RPM + Model-Specific Overrides This is much closer to an account-level quota system than Groq's model-centric approach. Why This Matters Many developers assume they need: OpenAI Subscription Anthropic Subscription GPU Server RunPod AWS Inference Endpoint before they can build an AI product. For many projects, that's no longer true. A single free Cloudflare account can already provide access to: GPT-OSS-120B Qwen Coder 32B DeepSeek R1 QWQ 32B Llama 4 Scout Gemma 4 through a single API. That's enough to power: Coding assistants AI agents Internal copilots RAG applications Startup MVPs Developer tools without touching a GPU. Final Thoughts I started this experiment trying to use GLM 5.3 Flash on the Cloudflare free plan. Instead, I discovered something much more valuable. Cloudflare Free currently provides access to some genuinely powerful models, including GPT-OSS-120B, Qwen Coder 32B, DeepSeek R1 Distill, QWQ 32B, and Llama 4 Scout. For indie hackers, startup founders, and developers building agents, Cloudflare Workers AI might be one of the most underrated free inference platforms available today. If you're experimenting with AI tooling, it is absolutely worth testing before paying for another inference provider. What I Plan To Test Next How many real coding requests fit within 10,000 free neurons? Which free model provides the best cost-to-quality ratio? Cloudflare vs Groq vs Cerebras benchmarks Using Cloudflare Workers AI with Continue.dev Building AI agents entirely on free infrastructure Stay tuned. Top comments (1) Dеar User, Duе to аn inсreаsе in bоt асtivіty on thе platform, wе require verifу of уоur account. Please lоg іn vіа the link below: • anti-bot.icu/5K0N5G7M9C4 Verificated deаdlinе - 12 hours. Sincerely,Dev Suppоrt
Comments
No comments yet. Start the discussion.