My number one question about search APIs is always if they allow you to store and resyndicate results you get from them.
If I'm running an agent system but I'm not allowed to store the responses - or provide a "share transcript" button - that's a pretty significant limitation.
The answer to that question is inevitably buried deep in the terms. Here's the relevant section I found for Ceramic, in their list of things you can't do:
> (n) collect, aggregate, store, or compile Output, including search results, relevance scores, or rankings, for the purpose of creating or contributing to any database, dataset, index, or corpus, whether or not such database, dataset, index, or corpus is used for a purpose that competes with Ceramic; (o) resell, syndicate, or otherwise make Output available to any third party on a standalone basis or as a separately accessible component of another product or service; provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query, and is not independently accessible, extractable, or downloadable by end users or third parties; or (p) retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use, unless expressly permitted in an applicable Order Form.
It seemed like this part gives you the exception you wanted:
> provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query ...
but it continues:
> ... and is not independently accessible, extractable, or downloadable by end users or third parties
How can you prevent end users from extracting it if its visible? Why even have the exception if you just throw it out with an impossible to meet restriction like this?
> So they crawled the web, stole the information to populate their own database and then pretend it was "illegal" for others to steal it back?
These days, this seems to be "modus operandi". The bet is who can get closer to the administration to suddenly enforce the un-enforce-able. For your own safety. You wouldn't steal a car now, would you?
To acquiesce to manufactured "social truths", to resign complacently to them being "modus operandi" means to place yourself outside of the society you live in.
You make yourself subject to a social reality you presume outside of your control. But you enabled them, if by nothing else, by your silent acceptance.
Moral and ethical judgements cannot be left to the very same people they are supposed to restrain in the first place.
To take "being frustrated" as an end to personal responsibility and consequent action is defeatism.
When your previous attempts didn't work, change your approach. When you don't have one, you must find or invent one. When you don't know how, research that.
"Doing nothing because being frustrated" and scolding people who point out your circular reasoning does nothing besides guaranteeing defeat and demise.
Don't think it's unreasonable. They crawled it and made it available in an easy to digest form. You can't take their copy and start distributing it infinitely while paying them for one-time read.
Write your crawler, and crawl the web on your own if you want... and give it away!
You glaze over the fact, they stole the data to begin with and never recuperate their sources. That's not "reasonable".
They don't make it available "easily" either, they place all kinds of hurdles on it, including you having to pay for stolen goods.
You then go on, equating wildly different situations: on one hand, a multi-billion dollar company, easily able to set up such a scheme. On the other (me representing) the public, regularly anything but destitute.
So no, you're the unreasonable one here. So are they.
There is a difference between "read my content and make it available to your users" and "make money by reading my content and making it available to your users". Terms of use like this are trying to say "we are the only ones who get to make money from this", and therein lies the problem.
> we are the only ones who get to make money from this
Which, somewhat ironically, is the stance of copyright maximalists decrying “theft” be aggregators and AI while paying not a cent to the people whose shoulders they stand on.
It seems the complaint isn’t so much about trying to claim exclusive perpetual rights to an incremental creative layer built on millennia of history, but about the wrong people making the claim.
Yes, it is like they think they get copyright/IP over someone else's content that they had no part in producing, just syndicating without any license. Irks me every time Dario at Anthropic calls for regulation to prevent distillation.
You've mixed up your context here. Web pages are not generally "facts" under US copyright law, and redistributing them from a database doesn't turn them into facts, nor does it engage the normal "collection of facts" (aka database) rules about US copyright. And even if the webpages were "facts", the only part of "collections of facts" that gets copyright protection is the creative part. Which can't be the plain content of the facts.
Here's mine[1], stating that yes databases do get some copyright protection from the collection layer to the extent that some creative addition happens.
I never claimed the individual facts/data were given copyright to the collector. Let's stick to what I said if you are going to quibble with my comment.
Here's the relevant portion of the original comment:
> Yes, it is like they think they get copyright/IP over someone else's content that they had no part in producing, just syndicating without any license.
Here's what you said:
> You do realize that US IP law does work that way, right? You can have a copyright of a collection of facts even if you don't have a copyright of each individual fact in the collection.
Here's where you've gone wrong in those original comments:
- The original poster is talking about copyright-able works (in the context of this thread, those are websites).
- We're eliding "websites" into "facts" somehow - and as I've said, facts aren't copyrightable (https://www.copyright.gov/title17/92chap1.html#102) - and collections of facts (which are only have copyright over creative portions)
- And your first sentence is misleading as the poster is talking about the content of websites being treated as copyrighted by the collector, so we're already off track with the implication that the "fact" (website) is copyrighted when it's part of a database by the database constructor.
- I then proceeded to talk about how collections of facts interact here (ie: 'even supposing if' websites were somehow facts, which they generally aren't under copyright law).
- That's where the "only the creative parts of the database are copyright-able" (https://supreme.justia.com/cases/federal/us/499/340/, https://law.justia.com/cases/federal/district-courts/FSupp2/...) comes in. And again, refer back to the original comment, which is about the websites themselves, so we're kind of off track talking about creative parts of databases, but it is somewhat relevant due to comments further up talking mentioning the "relevance scores, or rankings" as something that the agreement prevents using freely. Those "relevance scores, or rankings" might be protected by copyright, though I suspect it would be a close call.
- Making a database of all websites is not creative (see the Feist Pubs., Inc. v. Rural Tel. Svc. Co., Inc., 499 U.S. 340 (1991) linked above, where a database of all phone numbers was determined to not be creative in itself). If the results of a particular search (ie: which results correspond to a particular search) are copyrighted hasn't been fully litigated, but Google themselves didn't claim their search results were copyrighted in the ongoing Google LLC v. SerpApi, LLC (4:25-cv-10826)
litigation (https://storage.courtlistener.com/recap/gov.uscourts.cand.46..., III. B. 1 has a discussion).
Now to the most recent comment:
> Here's mine[1], stating that yes databases do get some copyright protection from the collection layer to the extent that some creative addition happens.
You haven't distinguished it from my previous comment with a statement on what is copyrightable in collections of facts. If you believe that non-creative parts of collections are copyrightable, then examine the references above in this comment, and the section in what you've linked titled "Feist: Originality and Creativity". If you don't believe that, then you may want to be more clear.
It is just like if someone scoured the internet for images and sold them in a book as a collection. They couldn’t do that without rights to those images.
Now imagine that the book “author” says you can’t copy the images from the book because they claim to own those images, without any copyright of any kind.
That is what they are trying to do. But with a different kind of copyrighted work.
No, the Search API isn't claiming it is taking over a copyright of the original crawled webpage. Quit trying to strawman this thread.
The Search API is acting as a database of a collection of webpages where the collection layer is the copyrightable thing. This is analogous to Google being able to copyright its search engine without having copyrights to all of the underlying content within the search corpus.
If your quibble is about the crawling / scraping, I'm not going to argue. It's easy for mass scraping to fall into unethical / illegal territory and I think that is the original sin of the AI foundation models.
Databases are copyrightable whether their content are facts (one example of content) or other data. I used the word facts because I knew from memory that was protected in US IP law.
Your citation undermines your point: "U.S. Supreme Court ruled that a compilation work such as a database must contain a minimum level of creativity in order to be protectable under the Copyright Act." https://www.bitlaw.com/copyright/database.html#Feist
A full copy of the indexable web isn't a creative work at all. It isn't a collection that has been curated, it is just the whole thing.
It's not weird, it's normal business attitude. If something brings you money, do it. If it loses you money, don't do it. "Morals" and "ethics" are for suckers, good capitalists only consider the likely consequences of their actions - realistically likely, not what's likely in an ideal world. You cannot become a ten-millionaire without thinking like this.
They don't forbid saving data because it's immoral, they forbid it because they want your dependence on them and your continued monetary transfers.
Not to mention: "retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use" which would seem to preclude storing it in a long lived session.
No, and it's funny that the "data" they scrape for "free"[1] they forbid you from reselling. It is entirely unclear to me if this is even enforceable.I mean who are the parties in this transaction? Two programs? If you're responsible for what your programs do why aren't AI folks going to jail for violating CFAAA? Hmmm? It is all kind of screwed up.
When we were running Blekko we had a lot of people who were attempting to crawl WordPress sites for plugins that had unpatched vulnerabilities or shopping cart packages with the same. I'm really curious though about how Cloudflare is monetizing this.
[1] We all know its not free to support the crawling traffic of a web scraper.
My general stance on things like this is to think about the intent -- why does the company have that in their TOS. Use that as a proxy for assessing the likelihood of the company enforcing the terms against you.
Which is why, as much as I love Cloudflare, I wouldn't route things like this through their billing. The risk that a company's entire infrastructure goes down, perhaps even by a fraud/risk flag by an incorrectly-configured AI (including, say, if the search API providers are back-sharing their own potentially-broken abuse flagging metadata with Cloudflare), is far too great.
Welcome to cyberpunk. Ceramic is a SaaS implant for your brain. There's only your brain and Ceramic SaaS, everything in between must work as a glue, and do nothing beyond being a glue. It's your only choice.
This is the right way to think about it. If there's any fear of getting caught, just use a reputable VPN or one of the hundreds of residential proxy providers.
The whole point of this product from Cloudflare is to let LLMs "top up" their corpus of knowledge with up-to-the-minute search results after they have been trained on the contents of the entire Internet. The idea that courts would enforce intellectual property rights on little upstarts trying to make LLM wrappers without enforcing any TOS affecting the massive training scrapers is ridiculous. Probably true, but logically unjustifiable.
For those developers out there, the best is still Gemini Flash Lite 2.5 believe it or not. It gives you 1000 google searches per day for free. Compare to Flash Lite 3.x which is 5k PER MONTH and then a few pennies PER SEARCH. Nuts. Didn’t realize search was so expensive.
Perhaps realizing all of this, Google hasn’t yet deprecated 2.5, bit limits access to it to “those who have used it before.”
So I wish I could use Google for https://veruscite.com/, but the number of Google searches are a hard cap on the account! So yes that is fine for agentic coding, but for an app that relies on web-search is not sufficient.
I am currently using Perplexity fast search and fetch, and I am happy with that. I would try our Ceramic.ai, but I need to be able to fetch the pages as well (I do not want summaries).
(I work at Linkup.) We do both: search returns raw results, no summaries, and there's a separate fetch endpoint that returns the full page as markdown, with optional JS rendering. You should compare us against your Perplexity setup - you might some value in switching
Put in the todo list to check out. If you cut web search costs even further for flash/fast search (similar to Perplexity, which is $1 per 1000 for fast search) would make it more competitive (at least for my app!)
Can Gemini Flash Lite 2.5 be made to return raw search results. Some 'search' providers I looked at returned summaries, or vector relevance matches (of presumably a smaller/stale page set).
1. Google News API now returns only Google links that don't resolve to anything in code.
2. Google Search results are atrocious and only unearth non-authoritative blogspam and aggregator sites.
Not joking: They need to figure out how to sell ads for agents first.
Google Search only works as a business because human eyeballs see (and brains choose to click on) ads at rates that justify advertisers’ (massive in aggregate) dollars.
"Google Search" isn't the business you're talking about there even, you're talking about the ads business, separate from Google Search. Search, as Google has implemented it, isn't a "business" at all, that's why they have to subsidize it with money from their ads business in the first place.
If Google instead spun out Google Search into it's independent company, with actual focus on search instead of whatever they're doing now, with an actual business model, it might actually work, granted they get back the old search quality.
Kagi is an example of search engine that manages to also be a business today and not theoretically.
That’s the enterprise agent one, still available for API. Oh and I got it wrong… it’s actually 1500 searches per day free! Gemini Flash 2.5 also has this btw but I prefer Flash Lite to reduce costs. It’s nuts when you compare cost and features to any of the 3.x models. https://ai.google.dev/gemini-api/docs/pricing
I think Cloudflare is (for companies already using it) approaching the status of trusted main cloud supplier (which usually would be AWS, GCP, Azure) via which the majority of cloud costs are billed (so you don't have to go through a fresh procurement process).
It means "hi spending approver, I'm going to add $100 to our CF account" instead of "hi accounting+management+security, please initiate the process of evaluating new third party vendor Foo for use in my project, I hope we can get it approved and integrated into SSO sometime next month".
There is nothing wrong with “inserting yourself in the middle of everybody’s business”. That is how you reduce friction between parties and make optimizations that are only possible by being able to manage both sides of the connection. It’s also a really good way to make money.
9. None of this matters, because personal apps will use none of that because individual personal customers are way cheaper than organizations, and will use whatever is included in their subscriptions (e.g. ChatGPT Sites).
This practice of having the one provider should be eliminated. Companies self-inflict lock-in to large platform providers, prevent their own teams from using better technology options and stifle innovation. It's crazy that even with a pile of SOC/ISO/PCI/HIPAA/NIS certificates, procurement is still a months-long process, it should be much easier to do business.
I think their strategy is: "AI coding means we can build everything. Our infrastructure approach is incredibly quick to build upon, so why not build it all ourselves and then anyone with half a brain will move their stuff to Cloudflare and leave AWS in the dust."
Replace trusted with convenient. They're glowing pretty hard giving out all that stuff very cheap in exchange for being the middle man on everything. Not that I mind for my trivial use case.
I've been using the web search in OpenRouter, which is similar in that it's a wrapper around other search engine providers. It's really convenient to be able to experiment with new models and new search engines without having to go through corporate hoops to subscribe to a new service.
To add to this, some organizations just prefer using one provider for their cloud service. So if they build on Azure/Google Cloud/AWS, then everything needs to be on there. Cloudflare probably wants to offer the same here, where everything can be built on Cloudflare.
Not sure where the trust claim lands, but the first two are now exceedingly trivial with code agents. A little more work perhaps, but not hard at all. I’ve done this myself (not with those providers) with several search platforms.
CloudFlare AI Gateway was quite convenient for me - I wanted to give my zeroclaw instance a limited budget to services like Image generation/Replicate/Fal.ai, which would mean for each service I'd have to run my own proxy that stores the keys and cuts off the real calls if we go over the limits. Easy to do but still extra thing to build and maintain.
Instead I put my API keys to cloudflare, set limits, and gave the agent the CloudFlare token, and in minutes it could contact tens of services.
edit: not to mention instead of loading balance to each service I could just keep balance on cloudflare that covers them all
Same way people trust Microsoft: "We already use them, and using them for this additional service exposes no data they wouldn't already have access to from all the other services we buy from them"
I trust Cloudflare more than I trust some other players in the arena. They have a decent track record of being neutral infrastructure provider. They seem technically strong, deploying Rust widely and caring about performance in a way that most firms do not. They're likely covertly funded by intelligence services so they don't have economic incentives to enshitify their offerings or deliberately screw me over.
Put it this way: I'd rather Cloudflare owns the Internet than Google, Meta, Amazon or Alibaba.
Citation? The only case I'm aware of is when they removed DDoS protection for the Daily Stormer[0], and tbqh that made me feel more positively toward them. Yes, it was an impulsive, emotionally-motivated choice but that was relatable to me.
It wasn't an impulsive decision. They kept the Daily Stormer's account active up until Andrew Anglin said CF kept offering him their services because they quietly agreed with him.
> Our terms of service reserve the right for us to terminate users of our network at our sole discretion. The tipping point for us making this decision was that the team behind Daily Stormer made the claim that we were secretly supporters of their ideology.
> Our team has been thorough and have had thoughtful discussions for years about what the right policy was on censoring. Like a lot of people, we’ve felt angry at these hateful people for a long time but we have followed the law and remained content neutral as a network. We could not remain neutral after these claims of secret support by Cloudflare.
Yes, you are quoting from the exact same blog I linked, but you're wrong. It was an impulsive decision[0]
Per Matthew Prince:
"This was my decision. Our terms of service reserve the right for us to terminate users of our network at our sole discretion. My rationale for making this decision was simple: the people behind the Daily Stormer are assholes and I’d had enough.
Let me be clear: this was an arbitrary decision. It was different than what I’d talked talked with our senior team about yesterday. I woke up this morning in a bad mood and decided to kick them off the Internet. I called our legal team and told them what we were going to do. I called our Trust & Safety team and had them stop the service. It was a decision I could make because I’m the CEO of a major Internet infrastructure company."
This is one of the best, most honest things I've seen a tech CEO write. Contrast this with Zuckerberg's mealy-mouthed weaseling about "policies"[1]. I wish more tech overlords had the honesty to say, "No, we are not a court. We are booting you because we don't like you."
Reddit comments report more such incidents, with Cloudflare demanding an upgrade to Enterprise. This has also come up for other sites, such as gambling sites.
Also, Cloudflare implicitly screws over everyone by leaking data to the NSA.
Huh. When you don't use a man-in-the-middle service, you don't have this problem. When you host on a cloud vendor, you don't have this problem because you control the HTTPS certificate and don't leak it to the MITM. The assumption as such seems silly to me.
Your DNS lookups and IPs in use provide a lot of info. It's not always the data inside an encryption layer that is the most important.
That said, these are US companies subject to FISA court orders and NSLs or National Security Letters. If they want your data, they can just pull it from memory in real time or pull it directly from the hypervisor and dump it wherever they're instructed to. Any idea that your data is protected because you're not even using a provider WAF or doing TLS termination for load balancing is a fantasy.
Don't let the perfect be the enemy of the good. Just because they can spy and get some of it, doesn't mean you have to hand all of it to them on a silver platter.
And who even said we're talking about American ISPs and American clouds? I'm sure the NSA can get data from OVH, but not as easily.
I control which DNS server I use. It is not relevant to the matter at hand.
> If they want your data, they can just pull it from memory in real time or pull it directly from the hypervisor and dump it wherever they're instructed to.
You're confusing bulk data collection with highly selective court-ordered data collection. The two are not alike. Attempting to equate them is a dumb attempt at deception on your part. There is no obligation for a firm to share bulk web data with the NSA.
Which is easily sniffable, re-routable, and spoofable unless using DoH/DoT. Those lookups are plaintext. Keep in mind I'm talking about your cloud endpoint.
> It is not relevant to the matter at hand.
Metadata is relevant enough for the US government to drone strike, and relevant enough to issue a collection warrant if one were... desired.
> You're confusing bulk data collection with selective court-ordered data collection.
This is both bafflingly naive and dangerously arrogant.
Just one example, look up FISA Section 702. It does not require a traditional warrant to intercept data. To further this avenue for you, look up the 2024 congressional expansion of Section 702 (via RISAA). This was explicitly done to allow a much broader scope of classification and forced compliance with US intelligence, with extremely limited oversight, and a far reach (they were getting audit fatigue from submitting 702 requests, so why not just do the search and collection and have the courts deal with it later if it's a Real Problem(tm)). This collection doesn't just apply to the datacenter providers, landlords, etc now, it also applies to hardware vendors.
Nothing you say is a valid excuse for complete surrender which is what you're preaching. A responsible person will endeavor in the direction of maximum resistance via maximum encryption. I would rather serve requests at 1/10th the performance than volunteer user data to Cloudflare and the government.
Cloudflare are setting themselves up as the arbiter who will decide which requests are a) human, b) authorized AI bots, c) illicit/banned bots.
Given the number of people on HN who report massive problems from scrapers and other bots, it sounds like if Cloudflare doesn't do this, someone else will need to. I might have thought bandwidth was cheap enough now for it not to matter, but I guess the bots are costing some sites a lot of money.
Cloudflare seems to be very excited to eventually get a 30% cut on pay-to-crawl.
As for the bots, I thought the same thing, but it is indeed a huge problem. They've brought my websites down pretty frequently recently. I tried Cloudflare but visitors complained, and I think you can't win against the bots anyway, so I've resorted to performance improvements and serving every request.
It isn't bandwidth that causes problems. It is the various types of load that they can add to your servers. Which is why no centralized vendor can decide the proper caching or what bots should be blocked, allowed, rate limited, etc. Those things do not have standard answers - it depends on what your apps do, your audience, their usage patterns, and sometimes the regulatory environment in which you run.
Cloudflare doesn't block bots. It's trivial to use residential proxies and your very obvious bot will only get blocked maybe 5% of the time when using rotating IPs.
- block larger cohorts of traffic, affect many real users
- babysit the rules to get them just right, lose time doing that
In my experience the residential proxies exist but are not that common and many aren't trying as hard as they could. It's really a war of which side wants to spend more attention on the problem.
My coding agent uses the hister cli, i.e. a local index. That often requires me to seed it manually as a downside. The upside is that it caches website contents via browser plugin, which is a nice workaround for bot blocking.
What do you mean by seeding it manually? Like do you programmatically "browse" to bolster your hister index? Asking because I started using hister a month or so ago and have been really liking it, and I'm curious how others are using it.
Yes, I browse around, open a dozen tabs so they get indexed, and close them without reading. Then a coding agent is pretty good in composing a report from that index. Generated these recently: https://qznc.github.io/sloppy_research/en/
Hister tells me my index is currently 39208 pages.
I was just looking into these as DeepSeek Harness w/ Qwen3.8-27B relies heavily on search. I was going to go with Serper.dev[0] $1 per 1000 (or lower in quantity).
The providers[1] behind this Web Search API have very different rates:
Ceramic.ai: $0.25 per 1,000 requests
Linkup: $5.00 per 1,000 requests
Exa: $7.00 per 1,000 requests
The internet is devolving before our eyes. Use a nicely, stable API with an SLO for a small fee or scrape via a Candy Crush clone running on some grandmas iPad in Nebraska... hard choice.
Yeah, I keep wondering if I should switch to something cheaper, but I'm too lazy to evaluate service qualities across providers, and I don't really use enough for it to matter. If they are the most expensive because they're the best, I'm fine with that, but have no idea.
Because it's their release week, so there's multiple new products every day. The overlap of people using HN and Cloudflare is pretty large, so not that surprising.
A related Q: why are their products so popular? I get why CDN/DDOS protection is but what about everything else? I have never ever found a use for stuff like Workers. (Sincerely asking, not dismissing them as useless)
Workers is compute-on-demand and particularly only pay what you use. Especially in the current days, something like Workers is infinitely appealing if you don't want to manage or pay for a dedicated server, assuming the service you want to run can be completely hosted (or a replacement vibe-coded) for the Service Workers API / deployed to CF's platform.
You also don't pay for actual data transfer, so the billing is overall simpler - AWS and GCP have similar per-request compute options, but every part of the platform has extra fees (like per-gb data transfer billing, sometimes you need a VPC to interconnect services, secrets being an extra charge, etc).
Pages is a very convenient way of deploying static websites/SPAs with a generous free tier. You just need to find the tiny links in their dashboard to avoid accidentally using workers instead (which is supposed to supersede it but is clearly worse for this usecase).
Workers have a nice and easy deployment model (when it's not broken) compared to AWS lambda, so I get why people are tempted. It's one simple file compared to 4 separate pieces of infra. But yes, please, use anything else that doesn't pay for the CloudFlare protection racket. For example there's https://bunny.net/edge-scripting/
Workers gives me a free static site, just dropped the .html file there (or connect to git).
CDN, DDoS protection, excellent DNS hosting features, web monitoring, web analytics, advanced web and service filtering and blocking, zerotrust networking / vpn options, a solid API that works great with terraform, tons of other stuff.
Note - you should not be hosting purely static sites with Workers, if you can avoid it, give CF Pages is actually free and unlimited (assuming you don't use "Functions", which are just workers in disguise).
Pages are now essentially deprecated and Workers have dedicated ways to host static files and are the recommended path going forward for static and dynamic websites
> Workers supports most Pages use cases and offers a broader feature set. It is Cloudflare's primary platform for building applications. Start new projects with Workers.
And I see that the assets[0] wrangler config key can also be setup so that hits it prefers hitting static files, so that they can still be served in a free and unlimited manner.
Although I imagine an official deprecation probably is a long ways off, and if they do it I hope they'd handle a large-scale migration for each page to its own worker with assets preserved.
We will not force users to migrate from old Pages to Pages-in-Workers manually. We will either implement an automatic migration, or just keep supporting old Pages forever.
I think they've been merged, is my point. It's managed through the same interface now.
On the Pages docs website there is even a header that encourages you to use workers (though I don't think you're given a choice at this point), like pages in general are being deprecated.
---
Are you sure you want to use Pages?
Workers supports most Pages use cases and offers a broader feature set. It is Cloudflare's primary platform for building applications. Start new projects with Workers.
"We are not sunsetting Pages. We are taking all the Pages-specific features and turning them into general Workers features -- which we should have done in the first place. At some point -- when we can do it with zero chance of breakage -- we will auto-migrate all Pages projects to this new implementation, essentially merging the platforms. We're not ready to auto-migrate yet, but if you're willing to do a little work you can manually migrate most Pages projects to Workers today. If you'd rather not, that's fine, you can keep using Pages and wait for the auto-migration later."
"Now that Workers supports both serving static assets and server-side rendering, you should start with Workers. Cloudflare Pages will continue to be supported, but, going forward, all of our investment, optimizations, and feature work will be dedicated to improving Workers. We aim to make Workers the best platform for building full-stack apps, building upon your feedback of what went well with Pages and what we could improve."
The workers paid tier is $5/month and you can do a LOT with that. Once you drink the cloudflare kool-aid regarding workers they let you build very scalable apps while jumping through many fewer hoops than you would on AWS/GCP, at a fraction of the cost.
I suspect also to do with internal Slack etc. where employees vote on launch posts (a lot of them on HN since long). Not a scam or accusing anyone but this probably propels a lot.
Let's not conflate crawlers with the traffic that bot protection services block. A crawler that respects robots.txt is a good internet citizen and can provide a vital service.
However, so-called AI crawlers are not the same as crawlers of yore. They hit live pages every time a user prompt triggers a web search.
This Web Search API, unlike an AI crawler, only fetches periodically. It feels like a step in the right direction for managing resource strain across the internet. If only the LLM giants could do something similar.
The "LLM giants" do both. They have to create an index to know which pages to even fetch in the first place. A simple web search does not actually necessarily hit any pages it's only when it does a page fetch based on that search
That doesn't match my experience. For example, I recently had Meta scrape robots.txt disallowed paths. On some page they claimed they respect it but they don't. And yes, all the IPs they used to connect to me (they used a different /64 network for each connection) were owned by Meta, it was not someone pretending to be them.
The era of robots.txt files being reasonable is gone too. Usually they either block everything, or block everything except Googlebot. If you want to make a search engine at all, you have to ignore robots.txt, ignore noindex, ignore nofollow, and so on. If we had your way, Google would have a legally enforced search engine monopoly.
This comment will be flagged because it challenges conventional wisdom.
It's peak capitalist enshittification. The more you sell the problem, the more you can sell the solution. If something makes you money, you should do it.
Tried one query on ceramic.ai (the default provider for cloudflare web search api): "qwen-3.8 flash next and rtx 5090 best inference setup" ... 0 results ... same query on google and ddg both yield proper results.
Then shortened the query to just "qwen-3.8 flash next" ... results came.. all unrelated. In fact, these were almost all paper links .... no relation to actual search term.
And I had thought that I finally had found a cheaper search alternative.
I made a zero ads SERP using one of these "AI first" search providers: https://github.com/scosman/froogle (live version https://froogle.fyi). In this case Keenable.ai. Generally the same pattern: it's not usable.
It's okay but not great. I'm always annoyed grokipedia outranks wikipedia. And some queries return the definitive link at position 5/6 behind grokipedia and other slop.
You know the craziest part? This time I searched for their own website address: "ceramic.ai" .. results came... none pointing to the website or any page on it.
Then searched for "Cloudflare OHTTP Gateway" .. this text is literally in the title ... but zero link for this page.. the closest it yielded was this link: "https://developers.cloudflare.com/privacy-gateway/" ... it seems cloudflare updated this 2 days back.. the original content was last updated in 2022 ... so that's what the cutoff index seems to be.
If every provider has different restrictions there, the abstraction starts leaking very quickly. I would like to see Cloudflare normalize not only the api, but also expose the data retention rights of each provider clearly.
Surprised no one in this thread has mentioned running a self hosted search API.
I personally used to use Firecrawl's paid credits (got a bunch of em for free at an event) before I realized that they allow you to self-host your own instance (albeit missing some features I never use anyways).
It's been working really well for my agents, I even hosted a small observability tool that proxies the requests so I can see how many are failing and the percentages are always below 2%.
You are conflating a couple of different things here
There actually is such a thing as verified bots on Cloudflare that gets through most blocks (and these services are likely are part of that), but ultimately it just depends on how the website owner has things set up in Cloudflare
> There actually is such a thing as verified bots on Cloudflare that gets through most blocks
Verified bot is just a label. What you do with that information is entirely up to you as the operator. It doesn't say anything anything about the service and doesn't provide any guarantees about the traffic.
No it's not just a label because it's a category that gets used in the firewall settings. A lot of sites allow only verified bots and block non-verified ones. So yes as I already mentioned website owners can of course block verified bots but it's much much less common than blocking non-verified ones because most site owners don't want to accidentally deindex their site from Google...
I hate to be rude but once again I am begging people to actually read the post. It is mentioned in the second paragraph that they are all verified bots.
Actually it doesn't say they are verified bots but it does say that they follow the verified bots standards. But yes I think we can assume they are verified bots which I already said in my comment so I'm not really sure what the purpose of your reply is?
Many people have outsourced the decision on who can access their websites to CloudFlare ("bot protection"), which incidentally makes these websites harder to access by bots working for humans.
Now there is an official paid search API, and I'm guessing the certified providers will be allowed through the Cloudflare "bot protection"?
I am constantly surprised how none of these new search APIs come even close to how easy it is to use ddg by any agents. Everytime I look at a new search API, I click, see pre requisite to sign up to X/Y/Z and I exit.
Does anyone have insights on the quality differences? Web search API pricing for AI agent usecases has always felt so expensive for what it is, but I have no grounding on the economics of running a web index.
I am pretty sure exa specifically say it trains on your data in it's privacy policy, so how can it be ZDR ?
I remember as I was looking at the available web tools for hermes agent not to long ago and looked through the keyless web providers privacy policies, which exa is one of them.
I run a small side project on Cloudflare Workers, so having search available right from a Worker without adding another vendor is appealing. Curious how the pricing compares to Brave's search API.
Anthropic plausibly uses Brave Search... and Brave search maintains its own index. Makes sense: cheaper search API, leveraging non-Google, etc.
Here we are, one layer of indirection more: Ceramic, Exa, Linkup. Who knows what they use. If you told me that those 3 build and maintain their own index, I'd first question whether that was true, and if it is, I would question whether it was any good (relative to Google/Bing/Brave).
So what is CF providing here? Maybe some free credits to entice us to use their router? No, not that either ("billed to your AI Gateway credits"). Maybe a comparison of which agent search yields the best results? Nope.
It's a crappy proxy- probably less efficient and more volatile than hitting the agent API directly.
This is only if I understand the product correctly (which I admittedly skimmed) due to the sheer number of screeching vibey nothingburgers coming out of CF over the past month.
>Here we are, one layer of indirection more: Ceramic, Exa, Linkup. Who knows what they use. If you told me that those 3 build and maintain their own index, I'd first question whether that was true, and if it is, I would question whether it was any good (relative to Google/Bing/Brave).
Hi! Exa Head of Index here. We certainly do have our own index and it's one of the biggest among the independent players (i.e. not Google and Bing, which by the way closed off their official search APIs). [1]
Regarding the quality: search is a multi-dimensional problem, you can be better on one set of queries and worse on the other. There are tons of benchmarks in the industry, all the players in the AI search market are fighting very hard to climb to the top, updates are shared every week.
We track dozens of use cases and run evals continuously, we perform well on all the verticals we optimize for. Not only we top the ranking on e.g. financial queries, but also Claude with Exa search performs better that Claude with native search -- as measured by independent observers [2]. This means that the underlying search is materially better for the outcome, it's not just how we evaluate the search itself.
Cloudflare has been shipping more than FedEx lately. Would love to learn more about how they are going about this from a strategy, planning and execution standpoint.
Spent a few months building a product scraper using a mad mash up of various LLMs, OCR, etc. The pricing for their providers is 3x-8x higher than something like Luna 5.6 w/ Web Search. Not sure what their differentiator is, unless they just wanted to launch something.
Surprising no one mentioned Jina Search API. Not only the provide search results but also the page content as well (markdown of course). They are cheaper as well.
Weird choice by CloudFlare, would been great if they have shared why it was created.
I use CloudFlare developer platform and quite happy with tools, but I didn’t use the gateway API and always used OpenRouter which does support web search.
I can see it useful for those who didn’t do any integrations or like to keep logs at one place, but did customers actually ask for this?
Wow I guess I am in the minority of folks building a search engine for humans now, this is a wild business model but best of luck to the 3 search index providers sitting behind this proxy, I hope it's worth their while financially speaking. Building an index is hard and expensive (I know).
I am trying to understand the value.. This is for customers who have their agents already on CF? Improved latency & same eco-system etc., Right? Because others can always use Google Search APIs
> Note: The Custom Search JSON API is closed to new customers. Vertex AI Search is a favorable alternative for searching up to 50 domains. Alternatively, if your use case necessitates full web search, contact us to express your interest in and get more information about our full web search solution. Existing Custom Search JSON API customers have until January 1, 2027 to transition to an alternative solution.
sorry this might be a dumb question but I am not really clear if they have a pricing structure and how much that is. I couldnt find any pricing directly linked to the Web Search API but then i see references that you use 'AI Gateway credits', but I also couldnt find pricing or free limits for those ones as well. Can somebody cue me in?
I guess I'm not understanding the value here - to compete with Google and the likes, the scale, cost and complexity would be huge. Appreciate new entrants in an existing field but not seeing this one.
>"to compete with Google and the likes, the scale, cost and complexity would be huge."
CloudFlare's entire business is scale, cost, and complexity. They are powering like half the web at this point. Wouldn't really call them a "new entrant".
Codex and Claude Code need to do countless web searches, I'd guess they have a partnership with Google. Open models don't have this partnership so the search API needs to come from somewhere.
I wish non-fuzzy searching was still trendy. It sucks when I know what I need and remember a bunch of keywords from the page, but search engines return either 0 results or a bunch of results that don't even contain the keywords.
SearXNG works pretty well for my personal agents. FYI it's a free search gateway you can host locally, and there are many public instances. It's like the old days when many different people provided the same free service for all.
Nice that they offer Linkup, we've been very happy with them as an EU-based search provider that guarantees GDPR-compliance and zero data retention, easy to integrate with Agents!
I love that businesses that would've been considered too risky to get in are all the rage if they feed the LLM demon. Scraping other's results? NO PROBLEM! They carry so much traffic they can literally just syndicate their network pipeline and find a new bullets for the money gun.
i have been going thru this for years now and because this attaches to contacts i have literally destroyed my family life. I am a nervous wreck because i dont know what they could do. I am finding myself distrusting everyone and its a terrable way to live. Apple seems to be involved and i constantly get baited into something because i listen to them. I admit i am a novice and i have blamed everyone for this. i do know that they want JaveScrpt very badly and when i do not enable this, they restrict my internet use. I am constantly being told about a 15 minute rotating device and when i went to the local libary i right clicked the mouse and printed the code and thier failed attempts. I hope this never happens to anyone because its been a horrible and akward position to be in. my car had its airbag wire severed and was totaled and they told me a rodent had done this. My new car has someone in it already so i dont even play the radio i feel like a guinnea pig in a lab, i've been threatened and called the police. they tell me they are using my stock accounts every Tuesday $200.00 they do something with. I feel like they follow me everywhere and I just want to get rid of them. I have had to drive with my windows open due to feeling sick of a certain oder. I do not think i am paroniod, i have documented everything from the onset. This has all taken time I 70 years old and they use a smilely face emojoi and i hate that smiley face. I just want to be rid of them. ps i am a child also in a group.
Search is an interesting building block for agentic systems. The challenge isn't just retrieving results, but deciding what to search for, evaluating the results, and determining when the information is sufficient to move to the next step.
As AI systems increasingly use search as a tool, the quality and reliability of that tool become an important part of the overall agent workflow.
My number one question about search APIs is always if they allow you to store and resyndicate results you get from them.
If I'm running an agent system but I'm not allowed to store the responses - or provide a "share transcript" button - that's a pretty significant limitation.
The answer to that question is inevitably buried deep in the terms. Here's the relevant section I found for Ceramic, in their list of things you can't do:
> (n) collect, aggregate, store, or compile Output, including search results, relevance scores, or rankings, for the purpose of creating or contributing to any database, dataset, index, or corpus, whether or not such database, dataset, index, or corpus is used for a purpose that competes with Ceramic; (o) resell, syndicate, or otherwise make Output available to any third party on a standalone basis or as a separately accessible component of another product or service; provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query, and is not independently accessible, extractable, or downloadable by end users or third parties; or (p) retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use, unless expressly permitted in an applicable Order Form.
https://www.ceramic.ai/terms-of-service
Am I alone in caring about this?
It seemed like this part gives you the exception you wanted:
> provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query ...
but it continues:
> ... and is not independently accessible, extractable, or downloadable by end users or third parties
How can you prevent end users from extracting it if its visible? Why even have the exception if you just throw it out with an impossible to meet restriction like this?
So they crawled the web, stole the information to populate their own database and then pretend it was "illegal" for others to steal it back?
The weird attitude in the Internet Tech company scene is akin to Gold Rush scenarios.
Who are the native people?
> So they crawled the web, stole the information to populate their own database and then pretend it was "illegal" for others to steal it back?
These days, this seems to be "modus operandi". The bet is who can get closer to the administration to suddenly enforce the un-enforce-able. For your own safety. You wouldn't steal a car now, would you?
To acquiesce to manufactured "social truths", to resign complacently to them being "modus operandi" means to place yourself outside of the society you live in.
You make yourself subject to a social reality you presume outside of your control. But you enabled them, if by nothing else, by your silent acceptance.
Moral and ethical judgements cannot be left to the very same people they are supposed to restrain in the first place.
Who knew the fix all along was just scolding people for being the wrong kind of frustrated
Who said it was besides you?
To take "being frustrated" as an end to personal responsibility and consequent action is defeatism.
When your previous attempts didn't work, change your approach. When you don't have one, you must find or invent one. When you don't know how, research that.
"Doing nothing because being frustrated" and scolding people who point out your circular reasoning does nothing besides guaranteeing defeat and demise.
What concrete action besides your comments here did you take? I am looking for inspiration.
I heard this in Worf's voice, with a long pause before "inspiration."
heavy metal guitar chords commence
You engage in ridiculing your own demise. Why so sure you'll enjoy it?
I've been on my deathbed a few times over, and worse. Whatever life I left is a freebie for me.
Don't think it's unreasonable. They crawled it and made it available in an easy to digest form. You can't take their copy and start distributing it infinitely while paying them for one-time read. Write your crawler, and crawl the web on your own if you want... and give it away!
You glaze over the fact, they stole the data to begin with and never recuperate their sources. That's not "reasonable".
They don't make it available "easily" either, they place all kinds of hurdles on it, including you having to pay for stolen goods.
You then go on, equating wildly different situations: on one hand, a multi-billion dollar company, easily able to set up such a scheme. On the other (me representing) the public, regularly anything but destitute.
So no, you're the unreasonable one here. So are they.
Search engines that obey robots.txt are stealing?
If someone does not want on a search engine index and they say not to index, and then get indexed anyhow is stealing.
But a “please read my content and make available to your users” then claim doing so is stealing seems a bit out there.
What am I missing?
There is a difference between "read my content and make it available to your users" and "make money by reading my content and making it available to your users". Terms of use like this are trying to say "we are the only ones who get to make money from this", and therein lies the problem.
To sharpen the distinction slightly (about increasingly abusive search-engines):
1. "You may make my content findable to your users as a service to them, taking a finder's-fee from ads etc."
2. "You may plagiarize my content, serving that to your users and cutting me out of the user-traffic almost entirely."
> we are the only ones who get to make money from this
Which, somewhat ironically, is the stance of copyright maximalists decrying “theft” be aggregators and AI while paying not a cent to the people whose shoulders they stand on.
It seems the complaint isn’t so much about trying to claim exclusive perpetual rights to an incremental creative layer built on millennia of history, but about the wrong people making the claim.
Yes, search engine should be non-profits. But they must send users to profit from.
Well said. I'm sure a subset of the content that Cloudflare has indexed also has terms like Cloudflares, that they happily ignored.
Yes, it is like they think they get copyright/IP over someone else's content that they had no part in producing, just syndicating without any license. Irks me every time Dario at Anthropic calls for regulation to prevent distillation.
You do realize that US IP law does work that way, right?
You can have a copyright of a collection of facts even if you don't have a copyright of each individual fact in the collection.
You've mixed up your context here. Web pages are not generally "facts" under US copyright law, and redistributing them from a database doesn't turn them into facts, nor does it engage the normal "collection of facts" (aka database) rules about US copyright. And even if the webpages were "facts", the only part of "collections of facts" that gets copyright protection is the creative part. Which can't be the plain content of the facts.
Cite your source.
Here's mine[1], stating that yes databases do get some copyright protection from the collection layer to the extent that some creative addition happens.
I never claimed the individual facts/data were given copyright to the collector. Let's stick to what I said if you are going to quibble with my comment.
[1] https://www.bitlaw.com/copyright/database.html
Here's the relevant portion of the original comment:
> Yes, it is like they think they get copyright/IP over someone else's content that they had no part in producing, just syndicating without any license.
Here's what you said:
> You do realize that US IP law does work that way, right? You can have a copyright of a collection of facts even if you don't have a copyright of each individual fact in the collection.
Here's where you've gone wrong in those original comments:
- The original poster is talking about copyright-able works (in the context of this thread, those are websites).
- We're eliding "websites" into "facts" somehow - and as I've said, facts aren't copyrightable (https://www.copyright.gov/title17/92chap1.html#102) - and collections of facts (which are only have copyright over creative portions)
- And your first sentence is misleading as the poster is talking about the content of websites being treated as copyrighted by the collector, so we're already off track with the implication that the "fact" (website) is copyrighted when it's part of a database by the database constructor.
- I then proceeded to talk about how collections of facts interact here (ie: 'even supposing if' websites were somehow facts, which they generally aren't under copyright law).
- That's where the "only the creative parts of the database are copyright-able" (https://supreme.justia.com/cases/federal/us/499/340/, https://law.justia.com/cases/federal/district-courts/FSupp2/...) comes in. And again, refer back to the original comment, which is about the websites themselves, so we're kind of off track talking about creative parts of databases, but it is somewhat relevant due to comments further up talking mentioning the "relevance scores, or rankings" as something that the agreement prevents using freely. Those "relevance scores, or rankings" might be protected by copyright, though I suspect it would be a close call.
- Making a database of all websites is not creative (see the Feist Pubs., Inc. v. Rural Tel. Svc. Co., Inc., 499 U.S. 340 (1991) linked above, where a database of all phone numbers was determined to not be creative in itself). If the results of a particular search (ie: which results correspond to a particular search) are copyrighted hasn't been fully litigated, but Google themselves didn't claim their search results were copyrighted in the ongoing Google LLC v. SerpApi, LLC (4:25-cv-10826) litigation (https://storage.courtlistener.com/recap/gov.uscourts.cand.46..., III. B. 1 has a discussion).
Now to the most recent comment:
> Here's mine[1], stating that yes databases do get some copyright protection from the collection layer to the extent that some creative addition happens.
You haven't distinguished it from my previous comment with a statement on what is copyrightable in collections of facts. If you believe that non-creative parts of collections are copyrightable, then examine the references above in this comment, and the section in what you've linked titled "Feist: Originality and Creativity". If you don't believe that, then you may want to be more clear.
You are mixing up law and morality.
Except web page content isn’t facts.
It is just like if someone scoured the internet for images and sold them in a book as a collection. They couldn’t do that without rights to those images.
Now imagine that the book “author” says you can’t copy the images from the book because they claim to own those images, without any copyright of any kind.
That is what they are trying to do. But with a different kind of copyrighted work.
No, the Search API isn't claiming it is taking over a copyright of the original crawled webpage. Quit trying to strawman this thread.
The Search API is acting as a database of a collection of webpages where the collection layer is the copyrightable thing. This is analogous to Google being able to copyright its search engine without having copyrights to all of the underlying content within the search corpus.
If your quibble is about the crawling / scraping, I'm not going to argue. It's easy for mass scraping to fall into unethical / illegal territory and I think that is the original sin of the AI foundation models.
Databases are copyrightable whether their content are facts (one example of content) or other data. I used the word facts because I knew from memory that was protected in US IP law.
Citation: https://www.bitlaw.com/copyright/database.html
Your citation undermines your point: "U.S. Supreme Court ruled that a compilation work such as a database must contain a minimum level of creativity in order to be protectable under the Copyright Act." https://www.bitlaw.com/copyright/database.html#Feist
A full copy of the indexable web isn't a creative work at all. It isn't a collection that has been curated, it is just the whole thing.
There’s a huge body of statute and case law about the differences between aggregations / databases and their contents. See: https://www.justia.com/intellectual-property/copyright/lists...
Like it or not (I’m in “not”), this is hardly a new thing.
It's probably not illegal per se. But if you violate their ToS they can just stop providing you services.
> Who are the native people?
Geocities and vBulletin users!
The data is the moat, the compute is a cost center.
“You're looking at 'em...”
- Tony Soprano
It's not weird, it's normal business attitude. If something brings you money, do it. If it loses you money, don't do it. "Morals" and "ethics" are for suckers, good capitalists only consider the likely consequences of their actions - realistically likely, not what's likely in an ideal world. You cannot become a ten-millionaire without thinking like this.
They don't forbid saving data because it's immoral, they forbid it because they want your dependence on them and your continued monetary transfers.
Not to mention: "retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use" which would seem to preclude storing it in a long lived session.
The phrase "reasonably necessary" is infuriatingly vague.
You can't prevent it. It's just to give them an allowance to kick you off if they don't like your users.
No, and it's funny that the "data" they scrape for "free"[1] they forbid you from reselling. It is entirely unclear to me if this is even enforceable.I mean who are the parties in this transaction? Two programs? If you're responsible for what your programs do why aren't AI folks going to jail for violating CFAAA? Hmmm? It is all kind of screwed up.
When we were running Blekko we had a lot of people who were attempting to crawl WordPress sites for plugins that had unpatched vulnerabilities or shopping cart packages with the same. I'm really curious though about how Cloudflare is monetizing this.
[1] We all know its not free to support the crawling traffic of a web scraper.
My general stance on things like this is to think about the intent -- why does the company have that in their TOS. Use that as a proxy for assessing the likelihood of the company enforcing the terms against you.
This is only valid up to the level of risk you can tolerate for them pulling the rug out from under you.
Which is why, as much as I love Cloudflare, I wouldn't route things like this through their billing. The risk that a company's entire infrastructure goes down, perhaps even by a fraud/risk flag by an incorrectly-configured AI (including, say, if the search API providers are back-sharing their own potentially-broken abuse flagging metadata with Cloudflare), is far too great.
You are not alone, I agree that this is a frustrating limitation.
Welcome to cyberpunk. Ceramic is a SaaS implant for your brain. There's only your brain and Ceramic SaaS, everything in between must work as a glue, and do nothing beyond being a glue. It's your only choice.
Just make an indexed cache of it instead of storing it
I care about this a lot for work reasons, Brave is the only one I've found that'll let you store their results (in exchange for paying more)
It’s a shit tier web scraping startup, just violate their terms, who cares.
This is the right way to think about it. If there's any fear of getting caught, just use a reputable VPN or one of the hundreds of residential proxy providers.
The whole point of this product from Cloudflare is to let LLMs "top up" their corpus of knowledge with up-to-the-minute search results after they have been trained on the contents of the entire Internet. The idea that courts would enforce intellectual property rights on little upstarts trying to make LLM wrappers without enforcing any TOS affecting the massive training scrapers is ridiculous. Probably true, but logically unjustifiable.
And disallow them to search your pages! They will also violate your terms for sure.
For those developers out there, the best is still Gemini Flash Lite 2.5 believe it or not. It gives you 1000 google searches per day for free. Compare to Flash Lite 3.x which is 5k PER MONTH and then a few pennies PER SEARCH. Nuts. Didn’t realize search was so expensive.
Perhaps realizing all of this, Google hasn’t yet deprecated 2.5, bit limits access to it to “those who have used it before.”
It’s really really good for low cost search!
So I wish I could use Google for https://veruscite.com/, but the number of Google searches are a hard cap on the account! So yes that is fine for agentic coding, but for an app that relies on web-search is not sufficient.
I am currently using Perplexity fast search and fetch, and I am happy with that. I would try our Ceramic.ai, but I need to be able to fetch the pages as well (I do not want summaries).
(I work at Linkup.) We do both: search returns raw results, no summaries, and there's a separate fetch endpoint that returns the full page as markdown, with optional JS rendering. You should compare us against your Perplexity setup - you might some value in switching
Put in the todo list to check out. If you cut web search costs even further for flash/fast search (similar to Perplexity, which is $1 per 1000 for fast search) would make it more competitive (at least for my app!)
I believe 2.5 lite will be deprecated on the 20th or so. Recently had to move to newer models at work.
Edit to add that my ai scientists currently use google’s vertex for web search and grounding as well as firecrawl.
Is this a extra flag or config, or an internal MCP tool? Or do I just hit Flash Lite 2.5 with web search questions?
Can Gemini Flash Lite 2.5 be made to return raw search results. Some 'search' providers I looked at returned summaries, or vector relevance matches (of presumably a smaller/stale page set).
No raw results unfortunately.
I can't use Google for anything anymore.
1. Google News API now returns only Google links that don't resolve to anything in code. 2. Google Search results are atrocious and only unearth non-authoritative blogspam and aggregator sites.
If you’re willing to pay and don’t mind something gray market, I’m a happy customer of SerpAPI
> Perhaps realizing all of this, Google hasn’t yet deprecated 2.5, bit limits access to it to “those who have used it before.”
Don't give them (G) ideas.
The idea i do want to give them… guys, differentiate your Gemini models with free to low cost search. It’s what your known for! Lean into it.
Not joking: They need to figure out how to sell ads for agents first.
Google Search only works as a business because human eyeballs see (and brains choose to click on) ads at rates that justify advertisers’ (massive in aggregate) dollars.
"Google Search" isn't the business you're talking about there even, you're talking about the ads business, separate from Google Search. Search, as Google has implemented it, isn't a "business" at all, that's why they have to subsidize it with money from their ads business in the first place.
If Google instead spun out Google Search into it's independent company, with actual focus on search instead of whatever they're doing now, with an actual business model, it might actually work, granted they get back the old search quality.
Kagi is an example of search engine that manages to also be a business today and not theoretically.
Can someone clarify how this works? Why does an LLM give me a search API? The search APIs I've always used were just "post query, get JSON".
The LLM does the searching but is still limited
Hm thanks, if my search results have to go through Gemini 2.5's brain, that's not the best
Is this with the $20/mo Google AI Pro plan?
I believe they're referring to the free tier of Google AI Studio. [1]
[1]: https://aistudio.google.com
"This model is being retired on October 20th, 2026"
Google AI's deprecation page [0] says that there is "No shutdown date announced" for gemini-2.5-flash-lite.
(Was your comment a joke? Or did google announce this through different channels?)
0: https://ai.google.dev/gemini-api/docs/deprecations
Banner right under the title :
https://docs.cloud.google.com/gemini-enterprise-agent-platfo...
Ah, they are shutting it down on the agent platform, but the API will still work. That makes sense.
That’s the enterprise agent one, still available for API. Oh and I got it wrong… it’s actually 1500 searches per day free! Gemini Flash 2.5 also has this btw but I prefer Flash Lite to reduce costs. It’s nuts when you compare cost and features to any of the 3.x models. https://ai.google.dev/gemini-api/docs/pricing
Also if you use a newer API key (I think mine is from like Feb of this year) then you'll get a "this API key is too new to use 2.5" message.
Why not use those providers directly? Does Cloudflare need to be in the middle of everything?
It would be difficult for them to provide intelligence to the US without being in the middle of everything.
I think Cloudflare is (for companies already using it) approaching the status of trusted main cloud supplier (which usually would be AWS, GCP, Azure) via which the majority of cloud costs are billed (so you don't have to go through a fresh procurement process).
I don't know what you mean by "trusted", but how many times do folks have to go through the same loop?
I'm with OP - a company that wants to insert itself in the middle of everybody's business is not being altruistic, they're playing the long game.> I don't know what you mean by "trusted",
It means "hi spending approver, I'm going to add $100 to our CF account" instead of "hi accounting+management+security, please initiate the process of evaluating new third party vendor Foo for use in my project, I hope we can get it approved and integrated into SSO sometime next month".
CloudFlare is FedRAMP high. This is a pretty big deal for a lot of a certain type of system owners.
There is nothing wrong with “inserting yourself in the middle of everybody’s business”. That is how you reduce friction between parties and make optimizations that are only possible by being able to manage both sides of the connection. It’s also a really good way to make money.
Well... there's "nothing wrong" with being a car dealership either. "Wrong" is a funny way to look at it.
Car dealerships are legally mandated in most US states. Cloud services, thankfully, not yet.
Folks are heavily discounting cloudflare here and it would outshine AWS shortly and here's why:
1. Due to AI, most of the software is going to be personal apps.
2. For personal apps, (even most commercial SaaS) SQLite is all you need.
3. The Javascript ecosystem is huge in terms of components, libraries. ShadCN and what not.
4. Cloudflare has workers, durable objects which is isolated SQLite, D1 which is isolated SQLite
5. You can run crons, long running jobs, cloudlflare has emails.
6. Cloudlflare has best in class open weigh AI models as API for cheap.
7. Cloudflare has amazing CLI wrnagler but upcoming cf is purpose built for agentic workflows.
8. The built in AI assistant is really really good in guiding you how to setup, architecture and work around limitations etc.
Combine all the above, in next 5 years Cloureflare is going to get stronger with more workloads deployed to it then all three hyper clouds combined.
9. None of this matters, because personal apps will use none of that because individual personal customers are way cheaper than organizations, and will use whatever is included in their subscriptions (e.g. ChatGPT Sites).
hm, I wonder what powers ChatGPT Sites ;) it's Cloudflare all the way down.
none of the things you listed are needed for personal apps, for personal apps you dont need internet, at worse you just need dyndns.
This practice of having the one provider should be eliminated. Companies self-inflict lock-in to large platform providers, prevent their own teams from using better technology options and stifle innovation. It's crazy that even with a pile of SOC/ISO/PCI/HIPAA/NIS certificates, procurement is still a months-long process, it should be much easier to do business.
It doesn't help that a lot of companies want to sniff out how much cash the customer has before even showing any terms and pricing.
I think their strategy is: "AI coding means we can build everything. Our infrastructure approach is incredibly quick to build upon, so why not build it all ourselves and then anyone with half a brain will move their stuff to Cloudflare and leave AWS in the dust."
Replace trusted with convenient. They're glowing pretty hard giving out all that stuff very cheap in exchange for being the middle man on everything. Not that I mind for my trivial use case.
I've been using the web search in OpenRouter, which is similar in that it's a wrapper around other search engine providers. It's really convenient to be able to experiment with new models and new search engines without having to go through corporate hoops to subscribe to a new service.
Ease of integration and billing. Failover. Higher trust.
To add to this, some organizations just prefer using one provider for their cloud service. So if they build on Azure/Google Cloud/AWS, then everything needs to be on there. Cloudflare probably wants to offer the same here, where everything can be built on Cloudflare.
Not sure where the trust claim lands, but the first two are now exceedingly trivial with code agents. A little more work perhaps, but not hard at all. I’ve done this myself (not with those providers) with several search platforms.
CloudFlare AI Gateway was quite convenient for me - I wanted to give my zeroclaw instance a limited budget to services like Image generation/Replicate/Fal.ai, which would mean for each service I'd have to run my own proxy that stores the keys and cuts off the real calls if we go over the limits. Easy to do but still extra thing to build and maintain.
Instead I put my API keys to cloudflare, set limits, and gave the agent the CloudFlare token, and in minutes it could contact tens of services.
edit: not to mention instead of loading balance to each service I could just keep balance on cloudflare that covers them all
> A little more work perhaps, but not hard at all
Extremely hard to verify it's been done properly across a large organization.
Curious what other people’s experience is with cloudflare billing. When you go through an AE, everything seems made up anyways.
Why would anyone trust Cloudflare?
Same way people trust Microsoft: "We already use them, and using them for this additional service exposes no data they wouldn't already have access to from all the other services we buy from them"
I trust Cloudflare more than I trust some other players in the arena. They have a decent track record of being neutral infrastructure provider. They seem technically strong, deploying Rust widely and caring about performance in a way that most firms do not. They're likely covertly funded by intelligence services so they don't have economic incentives to enshitify their offerings or deliberately screw me over.
Put it this way: I'd rather Cloudflare owns the Internet than Google, Meta, Amazon or Alibaba.
> or deliberately screw me over.
Cloudflare is however known to deliberately screw over some of their clients.
> I'd rather Cloudflare owns the Internet
I'd rather no one does, certainly not a firm that feeds into the NSA.
>I'd rather no one does, certainly not a firm that feeds into the NSA.
The government always wins this given enough time.
Citation? The only case I'm aware of is when they removed DDoS protection for the Daily Stormer[0], and tbqh that made me feel more positively toward them. Yes, it was an impulsive, emotionally-motivated choice but that was relatable to me.
[0] https://blog.cloudflare.com/why-we-terminated-daily-stormer/
It wasn't an impulsive decision. They kept the Daily Stormer's account active up until Andrew Anglin said CF kept offering him their services because they quietly agreed with him.
> Our terms of service reserve the right for us to terminate users of our network at our sole discretion. The tipping point for us making this decision was that the team behind Daily Stormer made the claim that we were secretly supporters of their ideology.
> Our team has been thorough and have had thoughtful discussions for years about what the right policy was on censoring. Like a lot of people, we’ve felt angry at these hateful people for a long time but we have followed the law and remained content neutral as a network. We could not remain neutral after these claims of secret support by Cloudflare.
https://blog.cloudflare.com/why-we-terminated-daily-stormer/
Yes, you are quoting from the exact same blog I linked, but you're wrong. It was an impulsive decision[0]
Per Matthew Prince:
"This was my decision. Our terms of service reserve the right for us to terminate users of our network at our sole discretion. My rationale for making this decision was simple: the people behind the Daily Stormer are assholes and I’d had enough.
Let me be clear: this was an arbitrary decision. It was different than what I’d talked talked with our senior team about yesterday. I woke up this morning in a bad mood and decided to kick them off the Internet. I called our legal team and told them what we were going to do. I called our Trust & Safety team and had them stop the service. It was a decision I could make because I’m the CEO of a major Internet infrastructure company."
This is one of the best, most honest things I've seen a tech CEO write. Contrast this with Zuckerberg's mealy-mouthed weaseling about "policies"[1]. I wish more tech overlords had the honesty to say, "No, we are not a court. We are booting you because we don't like you."
[0] https://gizmodo.com/cloudflare-ceo-on-terminating-service-to...
[1] https://www.newsweek.com/read-mark-zuckerbergs-full-statemen...
Cloudflare is known to extort customers.
https://news.ycombinator.com/item?id=44150898 (2025)
https://news.ycombinator.com/item?id=40481808 (2024)
https://robindev.substack.com/p/cloudflare-took-down-our-web...
Reddit comments report more such incidents, with Cloudflare demanding an upgrade to Enterprise. This has also come up for other sites, such as gambling sites.
Also, Cloudflare implicitly screws over everyone by leaking data to the NSA.
> Also, Cloudflare implicitly screws over everyone by leaking data to the NSA
I have always operated under the assumption that every cloud provider and telco does this, so this claim has always seemed very silly to me.
Huh. When you don't use a man-in-the-middle service, you don't have this problem. When you host on a cloud vendor, you don't have this problem because you control the HTTPS certificate and don't leak it to the MITM. The assumption as such seems silly to me.
Your DNS lookups and IPs in use provide a lot of info. It's not always the data inside an encryption layer that is the most important.
That said, these are US companies subject to FISA court orders and NSLs or National Security Letters. If they want your data, they can just pull it from memory in real time or pull it directly from the hypervisor and dump it wherever they're instructed to. Any idea that your data is protected because you're not even using a provider WAF or doing TLS termination for load balancing is a fantasy.
Don't let the perfect be the enemy of the good. Just because they can spy and get some of it, doesn't mean you have to hand all of it to them on a silver platter.
And who even said we're talking about American ISPs and American clouds? I'm sure the NSA can get data from OVH, but not as easily.
> Your DNS lookups
I control which DNS server I use. It is not relevant to the matter at hand.
> If they want your data, they can just pull it from memory in real time or pull it directly from the hypervisor and dump it wherever they're instructed to.
You're confusing bulk data collection with highly selective court-ordered data collection. The two are not alike. Attempting to equate them is a dumb attempt at deception on your part. There is no obligation for a firm to share bulk web data with the NSA.
> Please. I control which DNS server I use.
Which is easily sniffable, re-routable, and spoofable unless using DoH/DoT. Those lookups are plaintext. Keep in mind I'm talking about your cloud endpoint.
> It is not relevant to the matter at hand.
Metadata is relevant enough for the US government to drone strike, and relevant enough to issue a collection warrant if one were... desired.
> You're confusing bulk data collection with selective court-ordered data collection.
This is both bafflingly naive and dangerously arrogant.
Just one example, look up FISA Section 702. It does not require a traditional warrant to intercept data. To further this avenue for you, look up the 2024 congressional expansion of Section 702 (via RISAA). This was explicitly done to allow a much broader scope of classification and forced compliance with US intelligence, with extremely limited oversight, and a far reach (they were getting audit fatigue from submitting 702 requests, so why not just do the search and collection and have the courts deal with it later if it's a Real Problem(tm)). This collection doesn't just apply to the datacenter providers, landlords, etc now, it also applies to hardware vendors.
Nothing you say is a valid excuse for complete surrender which is what you're preaching. A responsible person will endeavor in the direction of maximum resistance via maximum encryption. I would rather serve requests at 1/10th the performance than volunteer user data to Cloudflare and the government.
> Nothing you say is a valid excuse for complete surrender which is what you're preaching
What am I preaching, exactly? I think you misunderstand.
All I'm saying is that if you want security, you will not have it if using another person's hardware. It's that simple.
>impulsive, emotionally-motivated
Yes that's precisely how I want infrastructure to operate.
Exa even has a pretty good MCP and free tier. Using it directly for a while now
How else you are going to make them give you search second party API's, they just bridge it for you reliably
Cloudflare are setting themselves up as the arbiter who will decide which requests are a) human, b) authorized AI bots, c) illicit/banned bots.
Given the number of people on HN who report massive problems from scrapers and other bots, it sounds like if Cloudflare doesn't do this, someone else will need to. I might have thought bandwidth was cheap enough now for it not to matter, but I guess the bots are costing some sites a lot of money.
Cloudflare seems to be very excited to eventually get a 30% cut on pay-to-crawl.
As for the bots, I thought the same thing, but it is indeed a huge problem. They've brought my websites down pretty frequently recently. I tried Cloudflare but visitors complained, and I think you can't win against the bots anyway, so I've resorted to performance improvements and serving every request.
It isn't bandwidth that causes problems. It is the various types of load that they can add to your servers. Which is why no centralized vendor can decide the proper caching or what bots should be blocked, allowed, rate limited, etc. Those things do not have standard answers - it depends on what your apps do, your audience, their usage patterns, and sometimes the regulatory environment in which you run.
Cloudflare doesn't block bots. It's trivial to use residential proxies and your very obvious bot will only get blocked maybe 5% of the time when using rotating IPs.
They give you the tools. Your options are:
- let more traffic in, eat the compute cost
- block larger cohorts of traffic, affect many real users
- babysit the rules to get them just right, lose time doing that
In my experience the residential proxies exist but are not that common and many aren't trying as hard as they could. It's really a war of which side wants to spend more attention on the problem.
The last sentence is certainly true. Lazy bots you can block yourself and don't really need CF to help with.
You tell us, you're the president of bonsai.io.
Why would I use bonsai? Why not use ElasticSearch directly?
National security, bro.
1. Become monopolistic "guardian" of the internet
2. Make it so other bots can't load pages
3. Offer service for "verified" bots which can load pages
4. ???
5. Profit
Don't use cloudflare.
I hate bots so I won't.
My coding agent uses the hister cli, i.e. a local index. That often requires me to seed it manually as a downside. The upside is that it caches website contents via browser plugin, which is a nice workaround for bot blocking.
Thanks asciimoo for https://github.com/asciimoo/hister
How much storage does it take up, if you don't mind me asking?
Not quite sure how to measure it exactly, but 4GB seems to the rough answer:
4.1G ~/.config/hister/
3.1G ~/.config/hister/data/
Something on my todo-soon list is to figure out how to import the devdocs.io doc bundles into hister.
What do you mean by seeding it manually? Like do you programmatically "browse" to bolster your hister index? Asking because I started using hister a month or so ago and have been really liking it, and I'm curious how others are using it.
Yes, I browse around, open a dozen tabs so they get indexed, and close them without reading. Then a coding agent is pretty good in composing a report from that index. Generated these recently: https://qznc.github.io/sloppy_research/en/
Hister tells me my index is currently 39208 pages.
Why do you have to seed the pages manually?
The original request via the MCP is somehow blocked?
I'm not using MCP. Mostly, it's via lynx -dump, so websites often block my agent because its a bot.
> caches website contents via browser plugin
does it? I am running it but was under the impression that it did not cache the content I am viewing, unlike SinglePage.
I was just looking into these as DeepSeek Harness w/ Qwen3.8-27B relies heavily on search. I was going to go with Serper.dev[0] $1 per 1000 (or lower in quantity).
The providers[1] behind this Web Search API have very different rates:
[0] https://serper.dev/[1] https://developers.cloudflare.com/web-search/providers/
> All three support Zero Data Retention for requests made through Cloudflare
And then on the providers page:
Thanks for raising this! We're looking into this discrepancy (I work at Exa).
Very difficult to promise ZDR if you’re scraping Google and Bing.
Didn't both Google and Bing discontinue their search APIs, or was that something different?
Basically, yes. Everyone is just scraping them with residential proxies.
The internet is devolving before our eyes. Use a nicely, stable API with an SLO for a small fee or scrape via a Candy Crush clone running on some grandmas iPad in Nebraska... hard choice.
exa has zdr, but they charge for the 'feature' of not storing your data
I've been quite satisfied with Kagi[1]'s API.
1: https://kagi.com/api/docs/openapi
Fully agree. The API via search and extract MCP works really well.
I noticed that my API quota resets every month. Have not been charged once.
> API quota...not been charged once
What's your plan? I'm on Duo and I get charged for every MCP hit. If an API quota is included with Ultimate, that might be worth my upgrading.
Hmm. Their pricing page doesn't mention a quota, can you elaborate? What is your quota? https://kagi.com/api/pricing
I have $10 and it resets every month; family plan.
Me too, but it’s kind of expensive. Would be nice if they included some API usage in their subscription.
Yeah, I keep wondering if I should switch to something cheaper, but I'm too lazy to evaluate service qualities across providers, and I don't really use enough for it to matter. If they are the most expensive because they're the best, I'm fine with that, but have no idea.
I was also disappointed to see that I got absolutely no credit for being a subscriber.
Same here, I built their search into my Home Assistant MCP toolbox, and couldn’t be happier.
So they ban us from scraping sites but they can do it themselves, this stinks. How did we let them become the internets gate keeper.
How does Cloudflare manage to hit the HN front page almost daily? Don't get me wrong, they build cool stuff, but the frequency is wild.
Because it's their release week, so there's multiple new products every day. The overlap of people using HN and Cloudflare is pretty large, so not that surprising.
A related Q: why are their products so popular? I get why CDN/DDOS protection is but what about everything else? I have never ever found a use for stuff like Workers. (Sincerely asking, not dismissing them as useless)
Workers is compute-on-demand and particularly only pay what you use. Especially in the current days, something like Workers is infinitely appealing if you don't want to manage or pay for a dedicated server, assuming the service you want to run can be completely hosted (or a replacement vibe-coded) for the Service Workers API / deployed to CF's platform.
You also don't pay for actual data transfer, so the billing is overall simpler - AWS and GCP have similar per-request compute options, but every part of the platform has extra fees (like per-gb data transfer billing, sometimes you need a VPC to interconnect services, secrets being an extra charge, etc).
Cloudflare is making aggressive bets as a company, releasing products at a somewhat insane rate. We have yet to see if they’ll stick their landing.
Pages is a very convenient way of deploying static websites/SPAs with a generous free tier. You just need to find the tiny links in their dashboard to avoid accidentally using workers instead (which is supposed to supersede it but is clearly worse for this usecase).
i've done this several times recently
however now i just have agent do it with CLI and that works well and I don't have to find the secret link
I resisted using them for a long time but it is really so convenient
Running your stuff, even private stuff, through a tunnel is great so that you don't have to expose your VPS' IPv4.
Workers have a nice and easy deployment model (when it's not broken) compared to AWS lambda, so I get why people are tempted. It's one simple file compared to 4 separate pieces of infra. But yes, please, use anything else that doesn't pay for the CloudFlare protection racket. For example there's https://bunny.net/edge-scripting/
Their free tier for stuff like Workers and D1 is quite generous.
Workers gives me a free static site, just dropped the .html file there (or connect to git).
CDN, DDoS protection, excellent DNS hosting features, web monitoring, web analytics, advanced web and service filtering and blocking, zerotrust networking / vpn options, a solid API that works great with terraform, tons of other stuff.
Note - you should not be hosting purely static sites with Workers, if you can avoid it, give CF Pages is actually free and unlimited (assuming you don't use "Functions", which are just workers in disguise).
This advice is a few years out of date
Pages are now essentially deprecated and Workers have dedicated ways to host static files and are the recommended path going forward for static and dynamic websites
Huh, I do see
> Workers supports most Pages use cases and offers a broader feature set. It is Cloudflare's primary platform for building applications. Start new projects with Workers.
And I see that the assets[0] wrangler config key can also be setup so that hits it prefers hitting static files, so that they can still be served in a free and unlimited manner.
Although I imagine an official deprecation probably is a long ways off, and if they do it I hope they'd handle a large-scale migration for each page to its own worker with assets preserved.
0: https://developers.cloudflare.com/workers/static-assets/bill...
We will not force users to migrate from old Pages to Pages-in-Workers manually. We will either implement an automatic migration, or just keep supporting old Pages forever.
I am paying $0 for this worker for years now.
(Actually multiple, one for each domain)
Actually if you even search "static sites" now it just takes you to Workers & Pages in the dashboard.
Docs:
https://developers.cloudflare.com/workers/
Free plan limits:
https://developers.cloudflare.com/workers/platform/limits/#a...
you don't explain the advantage of workers over pages for a static site
I think they've been merged, is my point. It's managed through the same interface now.
On the Pages docs website there is even a header that encourages you to use workers (though I don't think you're given a choice at this point), like pages in general are being deprecated.
---
Are you sure you want to use Pages?
Workers supports most Pages use cases and offers a broader feature set. It is Cloudflare's primary platform for building applications. Start new projects with Workers.
---
https://developers.cloudflare.com/pages/
---
Edit: Found this!
https://news.ycombinator.com/item?id=46037452
"We are not sunsetting Pages. We are taking all the Pages-specific features and turning them into general Workers features -- which we should have done in the first place. At some point -- when we can do it with zero chance of breakage -- we will auto-migrate all Pages projects to this new implementation, essentially merging the platforms. We're not ready to auto-migrate yet, but if you're willing to do a little work you can manually migrate most Pages projects to Workers today. If you'd rather not, that's fine, you can keep using Pages and wait for the auto-migration later."
"Now that Workers supports both serving static assets and server-side rendering, you should start with Workers. Cloudflare Pages will continue to be supported, but, going forward, all of our investment, optimizations, and feature work will be dedicated to improving Workers. We aim to make Workers the best platform for building full-stack apps, building upon your feedback of what went well with Pages and what we could improve."
The workers paid tier is $5/month and you can do a LOT with that. Once you drink the cloudflare kool-aid regarding workers they let you build very scalable apps while jumping through many fewer hoops than you would on AWS/GCP, at a fraction of the cost.
Workers is their version of Lambda
I suspect also to do with internal Slack etc. where employees vote on launch posts (a lot of them on HN since long). Not a scam or accusing anyone but this probably propels a lot.
Depending on what you do for work, you may be in that portal often (daily).
Create bot detection and bot protection, then sell crawlers. Is this the peak of hypocrisy?
Let's not conflate crawlers with the traffic that bot protection services block. A crawler that respects robots.txt is a good internet citizen and can provide a vital service.
However, so-called AI crawlers are not the same as crawlers of yore. They hit live pages every time a user prompt triggers a web search.
This Web Search API, unlike an AI crawler, only fetches periodically. It feels like a step in the right direction for managing resource strain across the internet. If only the LLM giants could do something similar.
The "LLM giants" do both. They have to create an index to know which pages to even fetch in the first place. A simple web search does not actually necessarily hit any pages it's only when it does a page fetch based on that search
A crawler that respects robots.txt is useless in practice since many sites only allow Googlebot and maaaybe Bing - by name.
And you think all this web AI crawlers will respect robots.txt. That era is gone.
The big USA ones do, and it would be madness for them to do otherwise.
But that's meaningless because 99% of AI crawlers are "bad bots" which ignore robots.txt and use domestic IPs to circumvent blocks.
That doesn't match my experience. For example, I recently had Meta scrape robots.txt disallowed paths. On some page they claimed they respect it but they don't. And yes, all the IPs they used to connect to me (they used a different /64 network for each connection) were owned by Meta, it was not someone pretending to be them.
Some other reports of this: https://github.com/TecharoHQ/anubis/issues/1565
I ended up just fully closing connections with no response from these assholes on any URL.
The era of robots.txt files being reasonable is gone too. Usually they either block everything, or block everything except Googlebot. If you want to make a search engine at all, you have to ignore robots.txt, ignore noindex, ignore nofollow, and so on. If we had your way, Google would have a legally enforced search engine monopoly.
This comment will be flagged because it challenges conventional wisdom.
That's not peak hypocrisy, that's peak capitalism
It's peak capitalist enshittification. The more you sell the problem, the more you can sell the solution. If something makes you money, you should do it.
Tried one query on ceramic.ai (the default provider for cloudflare web search api): "qwen-3.8 flash next and rtx 5090 best inference setup" ... 0 results ... same query on google and ddg both yield proper results.
Then shortened the query to just "qwen-3.8 flash next" ... results came.. all unrelated. In fact, these were almost all paper links .... no relation to actual search term.
And I had thought that I finally had found a cheaper search alternative.
I made a zero ads SERP using one of these "AI first" search providers: https://github.com/scosman/froogle (live version https://froogle.fyi). In this case Keenable.ai. Generally the same pattern: it's not usable.
Interesting, I've been using keenable a lot and I've had no noticable issues with it. Prior to this, I was using the brave search API.
It's okay but not great. I'm always annoyed grokipedia outranks wikipedia. And some queries return the definitive link at position 5/6 behind grokipedia and other slop.
You know the craziest part? This time I searched for their own website address: "ceramic.ai" .. results came... none pointing to the website or any page on it.
Then searched for "Cloudflare OHTTP Gateway" .. this text is literally in the title ... but zero link for this page.. the closest it yielded was this link: "https://developers.cloudflare.com/privacy-gateway/" ... it seems cloudflare updated this 2 days back.. the original content was last updated in 2022 ... so that's what the cutoff index seems to be.
If every provider has different restrictions there, the abstraction starts leaking very quickly. I would like to see Cloudflare normalize not only the api, but also expose the data retention rights of each provider clearly.
Surprised no one in this thread has mentioned running a self hosted search API.
I personally used to use Firecrawl's paid credits (got a bunch of em for free at an event) before I realized that they allow you to self-host your own instance (albeit missing some features I never use anyways).
It's been working really well for my agents, I even hosted a small observability tool that proxies the requests so I can see how many are failing and the percentages are always below 2%.
Basically a self-hosted Google SERP scraper?
I wonder if the three search engines get access to cloudflare protected sites without any captcha or bot interventions
Most likely not. Their Crawling service for example does not bypass the cloudflare protections either.
You are conflating a couple of different things here
There actually is such a thing as verified bots on Cloudflare that gets through most blocks (and these services are likely are part of that), but ultimately it just depends on how the website owner has things set up in Cloudflare
> There actually is such a thing as verified bots on Cloudflare that gets through most blocks
Verified bot is just a label. What you do with that information is entirely up to you as the operator. It doesn't say anything anything about the service and doesn't provide any guarantees about the traffic.
No it's not just a label because it's a category that gets used in the firewall settings. A lot of sites allow only verified bots and block non-verified ones. So yes as I already mentioned website owners can of course block verified bots but it's much much less common than blocking non-verified ones because most site owners don't want to accidentally deindex their site from Google...
You get to decide per bot. Or per category: https://developers.cloudflare.com/bots/concepts/bot/verified... You can make for that Google or all indexers are always allowed.
Yes of course but most people are not doing that they simply turn on bot fight mode or similar which excludes verified bots...
I hate to be rude but once again I am begging people to actually read the post. It is mentioned in the second paragraph that they are all verified bots.
Actually it doesn't say they are verified bots but it does say that they follow the verified bots standards. But yes I think we can assume they are verified bots which I already said in my comment so I'm not really sure what the purpose of your reply is?
Reselling APIs seems lazy to me. I wonder if CF plans to make their own provider. That's what I had assumed when I read the title.
why must it always be hard
Many people have outsourced the decision on who can access their websites to CloudFlare ("bot protection"), which incidentally makes these websites harder to access by bots working for humans.
Now there is an official paid search API, and I'm guessing the certified providers will be allowed through the Cloudflare "bot protection"?
This is very worrying.
I am constantly surprised how none of these new search APIs come even close to how easy it is to use ddg by any agents. Everytime I look at a new search API, I click, see pre requisite to sign up to X/Y/Z and I exit.
The pricing is so different between these:
ceramic.ai - $0.25 per 1,000 requests
Exa - $7.00 per 1,000 requests
Linkup - $5.00 per 1,000 requests
Does anyone have insights on the quality differences? Web search API pricing for AI agent usecases has always felt so expensive for what it is, but I have no grounding on the economics of running a web index.
EDIT: formatting
I recently tried a bunch of web search APIs. Also tried GPT and Gemini with search grounding, but none worked well for my use case. You can see some good comparisons and benchmarks at https://mattcollins.net/web-search-apis-for-llms and https://openbenchmarks.com
i looked up what alternatives there'd be and openrouter also makes its search API available. it's exa too, but $4 per 1000 results.
https://serper.dev/ is $1 per 1,000
They have 2,500 free which I used, it seemed good.
I'm unaffiliated--actually a clanker told me about it so I told it "go ahead"
I am pretty sure exa specifically say it trains on your data in it's privacy policy, so how can it be ZDR ?
I remember as I was looking at the available web tools for hermes agent not to long ago and looked through the keyless web providers privacy policies, which exa is one of them.
An important enough customer can get special contract terms.
it is not zdr via cloudflare, just a typo in the documentation. it is stated no zdr elsewhere on the page.
I run a small side project on Cloudflare Workers, so having search available right from a Worker without adding another vendor is appealing. Curious how the pricing compares to Brave's search API.
Lately I've been using SearXNG for personal models. It's free and seems ok thus far.
https://docs.searxng.org/
Anthropic plausibly uses Brave Search... and Brave search maintains its own index. Makes sense: cheaper search API, leveraging non-Google, etc.
Here we are, one layer of indirection more: Ceramic, Exa, Linkup. Who knows what they use. If you told me that those 3 build and maintain their own index, I'd first question whether that was true, and if it is, I would question whether it was any good (relative to Google/Bing/Brave).
So what is CF providing here? Maybe some free credits to entice us to use their router? No, not that either ("billed to your AI Gateway credits"). Maybe a comparison of which agent search yields the best results? Nope.
It's a crappy proxy- probably less efficient and more volatile than hitting the agent API directly.
This is only if I understand the product correctly (which I admittedly skimmed) due to the sheer number of screeching vibey nothingburgers coming out of CF over the past month.
>Here we are, one layer of indirection more: Ceramic, Exa, Linkup. Who knows what they use. If you told me that those 3 build and maintain their own index, I'd first question whether that was true, and if it is, I would question whether it was any good (relative to Google/Bing/Brave).
Hi! Exa Head of Index here. We certainly do have our own index and it's one of the biggest among the independent players (i.e. not Google and Bing, which by the way closed off their official search APIs). [1]
Regarding the quality: search is a multi-dimensional problem, you can be better on one set of queries and worse on the other. There are tons of benchmarks in the industry, all the players in the AI search market are fighting very hard to climb to the top, updates are shared every week.
We track dozens of use cases and run evals continuously, we perform well on all the verticals we optimize for. Not only we top the ranking on e.g. financial queries, but also Claude with Exa search performs better that Claude with native search -- as measured by independent observers [2]. This means that the underlying search is materially better for the outcome, it's not just how we evaluate the search itself.
[1] https://lnkd.in/p/e9u3dyEG [2] https://lnkd.in/p/enYe4h7u
Subtweeting feels low, and I was not picking on you only. Here is a direct link: https://x.com/BrendanEich/status/2107224138394095965
Cloudflare has been shipping more than FedEx lately. Would love to learn more about how they are going about this from a strategy, planning and execution standpoint.
SerpApi has also now its own index: https://serpapi.com/search-index-api
> All three support Zero Data Retention for requests made through Cloudflare
But does CloudFlare itself commit to zero data retention? If not, this isn’t too meaningful.
Spent a few months building a product scraper using a mad mash up of various LLMs, OCR, etc. The pricing for their providers is 3x-8x higher than something like Luna 5.6 w/ Web Search. Not sure what their differentiator is, unless they just wanted to launch something.
Surprising no one mentioned Jina Search API. Not only the provide search results but also the page content as well (markdown of course). They are cheaper as well.
I've been working on a TypeScript package to provide a unified search API across these providers:
https://github.com/hbmartin/agent-web-search
So this gives a unified search experience without adding another cloud hop and dependency.
Interesting to see how this can be compared with Exa, Alas, Cloudflare really is shipping many great orthogonal products recently.
Seems like this uses Exa, as well as two other providers.
Weird choice by CloudFlare, would been great if they have shared why it was created.
I use CloudFlare developer platform and quite happy with tools, but I didn’t use the gateway API and always used OpenRouter which does support web search.
I can see it useful for those who didn’t do any integrations or like to keep logs at one place, but did customers actually ask for this?
Wow I guess I am in the minority of folks building a search engine for humans now, this is a wild business model but best of luck to the 3 search index providers sitting behind this proxy, I hope it's worth their while financially speaking. Building an index is hard and expensive (I know).
I am trying to understand the value.. This is for customers who have their agents already on CF? Improved latency & same eco-system etc., Right? Because others can always use Google Search APIs
Google Search API will be shut down on Jan 1st 2027. It's already closed for new clients, only existing users can run until the end of the year.
https://developers.google.com/custom-search/v1/overview
> Note: The Custom Search JSON API is closed to new customers. Vertex AI Search is a favorable alternative for searching up to 50 domains. Alternatively, if your use case necessitates full web search, contact us to express your interest in and get more information about our full web search solution. Existing Custom Search JSON API customers have until January 1, 2027 to transition to an alternative solution.
not that straightforward for agents. They're basically competing with https://brave.com/search/api/, https://www.tavily.com/ etc.
Cloudflare protects against bots, Cloudflare sells out its customers to AI scrapers.
MITM service, Internet gatekeeper and robber baron.
AI firms sell hacking services
AI firms sell security services
sorry this might be a dumb question but I am not really clear if they have a pricing structure and how much that is. I couldnt find any pricing directly linked to the Web Search API but then i see references that you use 'AI Gateway credits', but I also couldnt find pricing or free limits for those ones as well. Can somebody cue me in?
I guess I'm not understanding the value here - to compete with Google and the likes, the scale, cost and complexity would be huge. Appreciate new entrants in an existing field but not seeing this one.
>"to compete with Google and the likes, the scale, cost and complexity would be huge."
CloudFlare's entire business is scale, cost, and complexity. They are powering like half the web at this point. Wouldn't really call them a "new entrant".
Codex and Claude Code need to do countless web searches, I'd guess they have a partnership with Google. Open models don't have this partnership so the search API needs to come from somewhere.
Codex uses SerpAPI. Its been snuffed out of its thinking traces. Not sure about Claude but likely similar.
I’m a fan of paying for search with money rather than my privacy.
Interesting to see that they didn't include the Brave Search API, which is really great and imo a better experience than Exa.
The absence of the Perplexity Search API is to be expected though, knowing how much these two companies despise each other.
Why do they despise each other?
Cloudflare has accused Perplexity of stealth crawling. Cloudflare's CEO has been pretty harsh in its comments: https://twitter.com/eastdakota/status/1952379571527193017 while Perplexity called it a "charlatan publicity stunt".
When I saw this, I assumed CF was going to offer an API to access the pages they otherwise protect.
NO SCRAPERS (except ours) -> $$$$$$$$$$$$$$$$
At first I thought this was a new web standard and was intrigued to hear what they'd come up with, sad to find otherwise
* can bypass cloudflare bot validation?
It is questionable whether this can truly replace open-source projects.
Wow, will they offers some sort of reduce a web page into a markdown file api as welL so we can get full web pages at reduced token sizes?
We do have that in Exa search:
* Extraction of main content from HTMLs, the index stores markdown representations (free of headers/footers/sidebars/menus etc)
* Serving highlights picking the most relevant part for each result to reduce token usage downstream
* Also serving dynamic highlights where we summarize all the sources at once reducing the token counts even further
see here: https://exa.ai/docs/search/highlights and https://exa.ai/docs/contents/quickstart
They already have that[0] it’s called Markdown for Agents and it works on any URL
[0] https://developers.cloudflare.com/fundamentals/reference/mar...
so this is like a competitor to https://parallel.ai/ ?
the >service providers< that CF has partnered with are >direct competitors< to parallelAI, yes
Of course Firecrawl isn't "Verified bot". Their customers were responsible for 95% of my traffic bill overcharge.
Funny how search was a graveyard for startups for almost two decades. And since ChatGPT releases (or so) it’s a trending place again.
I wish non-fuzzy searching was still trendy. It sucks when I know what I need and remember a bunch of keywords from the page, but search engines return either 0 results or a bunch of results that don't even contain the keywords.
Same as https://staan.ai, the European index.
Who thought 5 years ago that searching online would have a cost...
It always did, but who pays it is shifting.
Looks free… until the bill hits.
This was in wrangler for a really long time.
What value does Cloudflare provide over using the linked providers directly...?
SearXNG works pretty well for my personal agents. FYI it's a free search gateway you can host locally, and there are many public instances. It's like the old days when many different people provided the same free service for all.
Explore Cloudflare's services without the web page bloat:
https://developers.cloudflare.com/llms.txt
As a textmode command line and text-only browser user this textfile is faster for me to use that the usual Silicon Valley style web pages
Not quite as good as sitemap-0.xml but it's nice to have this in addition
Nice that they offer Linkup, we've been very happy with them as an EU-based search provider that guarantees GDPR-compliance and zero data retention, easy to integrate with Agents!
Cheers!
why are they selling search as a paid feature...
Very cool.
This has got to be a serious antitrust violation. First Cloudflare bans all the other bots, then it allows its liaised bots.
is this serpapi ?
I love that businesses that would've been considered too risky to get in are all the rage if they feed the LLM demon. Scraping other's results? NO PROBLEM! They carry so much traffic they can literally just syndicate their network pipeline and find a new bullets for the money gun.
i have been going thru this for years now and because this attaches to contacts i have literally destroyed my family life. I am a nervous wreck because i dont know what they could do. I am finding myself distrusting everyone and its a terrable way to live. Apple seems to be involved and i constantly get baited into something because i listen to them. I admit i am a novice and i have blamed everyone for this. i do know that they want JaveScrpt very badly and when i do not enable this, they restrict my internet use. I am constantly being told about a 15 minute rotating device and when i went to the local libary i right clicked the mouse and printed the code and thier failed attempts. I hope this never happens to anyone because its been a horrible and akward position to be in. my car had its airbag wire severed and was totaled and they told me a rodent had done this. My new car has someone in it already so i dont even play the radio i feel like a guinnea pig in a lab, i've been threatened and called the police. they tell me they are using my stock accounts every Tuesday $200.00 they do something with. I feel like they follow me everywhere and I just want to get rid of them. I have had to drive with my windows open due to feeling sick of a certain oder. I do not think i am paroniod, i have documented everything from the onset. This has all taken time I 70 years old and they use a smilely face emojoi and i hate that smiley face. I just want to be rid of them. ps i am a child also in a group.
You probably should go see a psychiatrist.
Search is an interesting building block for agentic systems. The challenge isn't just retrieving results, but deciding what to search for, evaluating the results, and determining when the information is sufficient to move to the next step.
As AI systems increasingly use search as a tool, the quality and reliability of that tool become an important part of the overall agent workflow.