Web Search API
developers.cloudflare.com429 points by tosh 11 hours ago
429 points by tosh 11 hours ago
My number one question about search APIs is always if they allow you to store and resyndicate results you get from them.
If I'm running an agent system but I'm not allowed to store the responses - or provide a "share transcript" button - that's a pretty significant limitation.
The answer to that question is inevitably buried deep in the terms. Here's the relevant section I found for Ceramic, in their list of things you can't do:
> (n) collect, aggregate, store, or compile Output, including search results, relevance scores, or rankings, for the purpose of creating or contributing to any database, dataset, index, or corpus, whether or not such database, dataset, index, or corpus is used for a purpose that competes with Ceramic; (o) resell, syndicate, or otherwise make Output available to any third party on a standalone basis or as a separately accessible component of another product or service; provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query, and is not independently accessible, extractable, or downloadable by end users or third parties; or (p) retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use, unless expressly permitted in an applicable Order Form.
https://www.ceramic.ai/terms-of-service
Am I alone in caring about this?
It seemed like this part gives you the exception you wanted:
> provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query ...
but it continues:
> ... and is not independently accessible, extractable, or downloadable by end users or third parties
How can you prevent end users from extracting it if its visible? Why even have the exception if you just throw it out with an impossible to meet restriction like this?
So they crawled the web, stole the information to populate their own database and then pretend it was "illegal" for others to steal it back?
The weird attitude in the Internet Tech company scene is akin to Gold Rush scenarios.
Who are the native people?
> So they crawled the web, stole the information to populate their own database and then pretend it was "illegal" for others to steal it back?
These days, this seems to be "modus operandi". The bet is who can get closer to the administration to suddenly enforce the un-enforce-able. For your own safety. You wouldn't steal a car now, would you?
To acquiesce to manufactured "social truths", to resign complacently to them being "modus operandi" means to place yourself outside of the society you live in.
You make yourself subject to a social reality you presume outside of your control. But you enabled them, if by nothing else, by your silent acceptance.
Moral and ethical judgements cannot be left to the very same people they are supposed to restrain in the first place.
Who knew the fix all along was just scolding people for being the wrong kind of frustrated
Don't think it's unreasonable. They crawled it and made it available in an easy to digest form. You can't take their copy and start distributing it infinitely while paying them for one-time read. Write your crawler, and crawl the web on your own if you want... and give it away!
You glaze over the fact, they stole the data to begin with and never recuperate their sources. That's not "reasonable".
They don't make it available "easily" either, they place all kinds of hurdles on it, including you having to pay for stolen goods.
You then go on, equating wildly different situations: on one hand, a multi-billion dollar company, easily able to set up such a scheme. On the other (me representing) the public, regularly anything but destitute.
So no, you're the unreasonable one here. So are they.
Search engines that obey robots.txt are stealing?
If someone does not want on a search engine index and they say not to index, and then get indexed anyhow is stealing.
But a “please read my content and make available to your users” then claim doing so is stealing seems a bit out there.
What am I missing?
There is a difference between "read my content and make it available to your users" and "make money by reading my content and making it available to your users". Terms of use like this are trying to say "we are the only ones who get to make money from this", and therein lies the problem.
Not to mention: "retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use" which would seem to preclude storing it in a long lived session.
My general stance on things like this is to think about the intent -- why does the company have that in their TOS. Use that as a proxy for assessing the likelihood of the company enforcing the terms against you.
This is only valid up to the level of risk you can tolerate for them pulling the rug out from under you.
Which is why, as much as I love Cloudflare, I wouldn't route things like this through their billing. The risk that a company's entire infrastructure goes down, perhaps even by a fraud/risk flag by an incorrectly-configured AI (including, say, if the search API providers are back-sharing their own potentially-broken abuse flagging metadata with Cloudflare), is far too great.
It’s a shit tier web scraping startup, just violate their terms, who cares.
This is the right way to think about it. If there's any fear of getting caught, just use a reputable VPN or one of the hundreds of residential proxy providers.
The whole point of this product from Cloudflare is to let LLMs "top up" their corpus of knowledge with up-to-the-minute search results after they have been trained on the contents of the entire Internet. The idea that courts would enforce intellectual property rights on little upstarts trying to make LLM wrappers without enforcing any TOS affecting the massive training scrapers is ridiculous. Probably true, but logically unjustifiable.
For those developers out there, the best is still Gemini Flash Lite 2.5 believe it or not. It gives you 1000 google searches per day for free. Compare to Flash Lite 3.x which is 5k PER MONTH and then a few pennies PER SEARCH. Nuts. Didn’t realize search was so expensive.
Perhaps realizing all of this, Google hasn’t yet deprecated 2.5, bit limits access to it to “those who have used it before.”
It’s really really good for low cost search!
So I wish I could use Google for https://veruscite.com/, but the number of Google searches are a hard cap on the account! So yes that is fine for agentic coding, but for an app that relies on web-search is not sufficient.
I am currently using Perplexity fast search and fetch, and I am happy with that. I would try our Ceramic.ai, but I need to be able to fetch the pages as well (I do not want summaries).
(I work at Linkup.) We do both: search returns raw results, no summaries, and there's a separate fetch endpoint that returns the full page as markdown, with optional JS rendering. You should compare us against your Perplexity setup - you might some value in switching
Put in the todo list to check out. If you cut web search costs even further for flash/fast search (similar to Perplexity, which is $1 per 1000 for fast search) would make it more competitive (at least for my app!)
Can Gemini Flash Lite 2.5 be made to return raw search results. Some 'search' providers I looked at returned summaries, or vector relevance matches (of presumably a smaller/stale page set).
I can't use Google for anything anymore.
1. Google News API now returns only Google links that don't resolve to anything in code. 2. Google Search results are atrocious and only unearth non-authoritative blogspam and aggregator sites.
> Perhaps realizing all of this, Google hasn’t yet deprecated 2.5, bit limits access to it to “those who have used it before.”
Don't give them (G) ideas.
The idea i do want to give them… guys, differentiate your Gemini models with free to low cost search. It’s what your known for! Lean into it.
Not joking: They need to figure out how to sell ads for agents first.
Google Search only works as a business because human eyeballs see (and brains choose to click on) ads at rates that justify advertisers’ (massive in aggregate) dollars.
"Google Search" isn't the business you're talking about there even, you're talking about the ads business, separate from Google Search. Search, as Google has implemented it, isn't a "business" at all, that's why they have to subsidize it with money from their ads business in the first place.
If Google instead spun out Google Search into it's independent company, with actual focus on search instead of whatever they're doing now, with an actual business model, it might actually work, granted they get back the old search quality.
Kagi is an example of search engine that manages to also be a business today and not theoretically.
Is this with the $20/mo Google AI Pro plan?
"This model is being retired on October 20th, 2026"
Google AI's deprecation page [0] says that there is "No shutdown date announced" for gemini-2.5-flash-lite.
(Was your comment a joke? Or did google announce this through different channels?)
Banner right under the title :
https://docs.cloud.google.com/gemini-enterprise-agent-platfo...
Ah, they are shutting it down on the agent platform, but the API will still work. That makes sense.
That’s the enterprise agent one, still available for API. Oh and I got it wrong… it’s actually 1500 searches per day free! Gemini Flash 2.5 also has this btw but I prefer Flash Lite to reduce costs. It’s nuts when you compare cost and features to any of the 3.x models. https://ai.google.dev/gemini-api/docs/pricing
Also if you use a newer API key (I think mine is from like Feb of this year) then you'll get a "this API key is too new to use 2.5" message.
Can someone clarify how this works? Why does an LLM give me a search API? The search APIs I've always used were just "post query, get JSON".
The LLM does the searching but is still limited
Hm thanks, if my search results have to go through Gemini 2.5's brain, that's not the best
Why not use those providers directly? Does Cloudflare need to be in the middle of everything?
It would be difficult for them to provide intelligence to the US without being in the middle of everything.
I think Cloudflare is (for companies already using it) approaching the status of trusted main cloud supplier (which usually would be AWS, GCP, Azure) via which the majority of cloud costs are billed (so you don't have to go through a fresh procurement process).
I don't know what you mean by "trusted", but how many times do folks have to go through the same loop?
- Company has great initial product
- Company gets popular
- Shareholders demand infinite growth
- Company becomes rent-seeker
- GOTO 10
I'm with OP - a company that wants to insert itself in the middle of everybody's business is not being altruistic, they're playing the long game.> I don't know what you mean by "trusted",
It means "hi spending approver, I'm going to add $100 to our CF account" instead of "hi accounting+management+security, please initiate the process of evaluating new third party vendor Foo for use in my project, I hope we can get it approved and integrated into SSO sometime next month".
CloudFlare is FedRAMP high. This is a pretty big deal for a lot of a certain type of system owners.
There is nothing wrong with “inserting yourself in the middle of everybody’s business”. That is how you reduce friction between parties and make optimizations that are only possible by being able to manage both sides of the connection. It’s also a really good way to make money.
Well... there's "nothing wrong" with being a car dealership either. "Wrong" is a funny way to look at it.
Replace trusted with convenient. They're glowing pretty hard giving out all that stuff very cheap in exchange for being the middle man on everything. Not that I mind for my trivial use case.
This practice of having the one provider should be eliminated. Companies self-inflict lock-in to large platform providers, prevent their own teams from using better technology options and stifle innovation. It's crazy that even with a pile of SOC/ISO/PCI/HIPAA/NIS certificates, procurement is still a months-long process, it should be much easier to do business.
It doesn't help that a lot of companies want to sniff out how much cash the customer has before even showing any terms and pricing.
I think their strategy is: "AI coding means we can build everything. Our infrastructure approach is incredibly quick to build upon, so why not build it all ourselves and then anyone with half a brain will move their stuff to Cloudflare and leave AWS in the dust."
I've been using the web search in OpenRouter, which is similar in that it's a wrapper around other search engine providers. It's really convenient to be able to experiment with new models and new search engines without having to go through corporate hoops to subscribe to a new service.
Ease of integration and billing. Failover. Higher trust.
To add to this, some organizations just prefer using one provider for their cloud service. So if they build on Azure/Google Cloud/AWS, then everything needs to be on there. Cloudflare probably wants to offer the same here, where everything can be built on Cloudflare.
Not sure where the trust claim lands, but the first two are now exceedingly trivial with code agents. A little more work perhaps, but not hard at all. I’ve done this myself (not with those providers) with several search platforms.
CloudFlare AI Gateway was quite convenient for me - I wanted to give my zeroclaw instance a limited budget to services like Image generation/Replicate/Fal.ai, which would mean for each service I'd have to run my own proxy that stores the keys and cuts off the real calls if we go over the limits. Easy to do but still extra thing to build and maintain.
Instead I put my API keys to cloudflare, set limits, and gave the agent the CloudFlare token, and in minutes it could contact tens of services.
edit: not to mention instead of loading balance to each service I could just keep balance on cloudflare that covers them all
> A little more work perhaps, but not hard at all
Extremely hard to verify it's been done properly across a large organization.
Curious what other people’s experience is with cloudflare billing. When you go through an AE, everything seems made up anyways.
Why would anyone trust Cloudflare?
Same way people trust Microsoft: "We already use them, and using them for this additional service exposes no data they wouldn't already have access to from all the other services we buy from them"
I trust Cloudflare more than I trust some other players in the arena. They have a decent track record of being neutral infrastructure provider. They seem technically strong, deploying Rust widely and caring about performance in a way that most firms do not. They're likely covertly funded by intelligence services so they don't have economic incentives to enshitify their offerings or deliberately screw me over.
Put it this way: I'd rather Cloudflare owns the Internet than Google, Meta, Amazon or Alibaba.
> or deliberately screw me over.
Cloudflare is however known to deliberately screw over some of their clients.
> I'd rather Cloudflare owns the Internet
I'd rather no one does, certainly not a firm that feeds into the NSA.
>I'd rather no one does, certainly not a firm that feeds into the NSA.
The government always wins this given enough time.
Citation? The only case I'm aware of is when they removed DDoS protection for the Daily Stormer[0], and tbqh that made me feel more positively toward them. Yes, it was an impulsive, emotionally-motivated choice but that was relatable to me.
[0] https://blog.cloudflare.com/why-we-terminated-daily-stormer/
It wasn't an impulsive decision. They kept the Daily Stormer's account active up until Andrew Anglin said CF kept offering him their services because they quietly agreed with him.
> Our terms of service reserve the right for us to terminate users of our network at our sole discretion. The tipping point for us making this decision was that the team behind Daily Stormer made the claim that we were secretly supporters of their ideology.
> Our team has been thorough and have had thoughtful discussions for years about what the right policy was on censoring. Like a lot of people, we’ve felt angry at these hateful people for a long time but we have followed the law and remained content neutral as a network. We could not remain neutral after these claims of secret support by Cloudflare.
https://blog.cloudflare.com/why-we-terminated-daily-stormer/
Yes, you are quoting from the exact same blog I linked, but you're wrong. It was an impulsive decision[0]
Per Matthew Prince:
"This was my decision. Our terms of service reserve the right for us to terminate users of our network at our sole discretion. My rationale for making this decision was simple: the people behind the Daily Stormer are assholes and I’d had enough.
Let me be clear: this was an arbitrary decision. It was different than what I’d talked talked with our senior team about yesterday. I woke up this morning in a bad mood and decided to kick them off the Internet. I called our legal team and told them what we were going to do. I called our Trust & Safety team and had them stop the service. It was a decision I could make because I’m the CEO of a major Internet infrastructure company."
This is one of the best, most honest things I've seen a tech CEO write. Contrast this with Zuckerberg's mealy-mouthed weaseling about "policies"[1]. I wish more tech overlords had the honesty to say, "No, we are not a court. We are booting you because we don't like you."
[0] https://gizmodo.com/cloudflare-ceo-on-terminating-service-to...
[1] https://www.newsweek.com/read-mark-zuckerbergs-full-statemen...
Cloudflare is known to extort customers.
https://news.ycombinator.com/item?id=44150898 (2025)
https://news.ycombinator.com/item?id=40481808 (2024)
https://robindev.substack.com/p/cloudflare-took-down-our-web...
Reddit comments report more such incidents, with Cloudflare demanding an upgrade to Enterprise. This has also come up for other sites, such as gambling sites.
Also, Cloudflare implicitly screws over everyone by leaking data to the NSA.
> Also, Cloudflare implicitly screws over everyone by leaking data to the NSA
I have always operated under the assumption that every cloud provider and telco does this, so this claim has always seemed very silly to me.
Huh. When you don't use a man-in-the-middle service, you don't have this problem. When you host on a cloud vendor, you don't have this problem because you control the HTTPS certificate and don't leak it to the MITM. The assumption as such seems silly to me.