Billing starts on 1 November
Cloudflare's AI Search stops being free on 1 November 2026, and from then the type of query a team runs sets most of its bill. Cloudflare made the managed retrieval service generally available on 1 October, during its Birthday Week, in a blog post.
AI Search, a managed retrieval pipeline built on Workers AI, Vectorize, R2 and Browser Run, was free in beta, according to Cloudflare's August preview post. Beyond the free allotment, the limits and pricing page lists ingestion at $0.75 per million tokens, $0.50 more for images, and storage at $2.00 per GB-month.
Semantic queries, which cover vector and hybrid search, cost $0.75 per 1,000, against $0.10 for full-text queries. That makes a semantic query 7.5 times the price of a full-text one. Beyond the allotment, 100,000 queries a month would cost $75 on semantic search and $10 on full-text.
The free query pool splits in two
Every Workers plan gets 5 million ingestion tokens, 10 GB of storage, 1,000 semantic queries and 1,000 full-text queries free each month. The August preview offered a single pool of 2,000 queries shared across both types. Every other rate and allotment in the GA table matches the preview table.
Cloudflare's preview post priced a sample instance with 30,000 semantic queries a month at $21.00 for queries, after 2,000 free ones. Under the GA allotment, the same queries would cost $21.75, and $0.75 a month is the most the split can add for any team.
Workers AI embedding and reranking are included in AI Search usage and do not appear on the Workers AI bill, the docs say. Generation, query rewriting and external model providers still bill through their own services. Cloudflare says it will send a reminder email the week before billing begins.
The same week, Cloudflare put a Web Search API into beta through AI Gateway, where searches bill at each provider's list price. See Cloudflare Web Search API: Exa costs 28 times Ceramic per search, and Cloudflare's retention pages disagree.
Image embeddings, OCR and larger files
- Native image retrieval works with the Qwen3-VL-Embedding model, the post says. Instances with a text-only model can still take image queries, which AI Search converts to a caption first.
- OCR reads scanned PDFs before chunking and bills as image processing ingestion tokens.
- The post says text files and PDFs can now run to 10 MiB, up from 4 MiB.
- The limits page is narrower on PDFs. A PDF reaches 10 MiB only with OCR enabled, while PDFs without OCR and other formats stay at 4 MiB.
- Tokens are counted with the cl100k_base tokenizer on the final chunks, so text repeated by chunk overlap is counted more than once.
Free plan crawls stop at 500 pages a day
- On Workers Free, website crawls stop at 500 pages a day, which the docs call the binding limit.
- An instance holds up to 100,000 files on Workers Free and 1 million on Workers Paid, or 500,000 with hybrid search.
- Neither the post nor the docs give retrieval quality or latency figures for the new image embeddings.
- Cloudflare says it is refactoring the keyword search engine, because the current one has limits with big data stores.
Check old R2 buckets before November
- Teams whose instances crawled websites should look for the R2 bucket AI Search originally created. The docs say AI Search no longer uses it, but objects left in it may still count toward R2 storage.
- Teams can estimate ingestion before 1 November with any cl100k_base tokenizer, such as tiktoken, counting chunk overlap.
- Teams that run semantic search at volume should price it at $0.75 per 1,000 queries, against $0.10 for full-text.
