According to Cloudflare, the amount of traffic generated by machines (AI agents, crawlers) on the internet has already exceeded that of humans. Their CEO originally estimated this would happen by the end of 2027, later changed it to early 2027, but in June 2026, he himself posted that it had already happened earlier than that.
First, clarify the scope: This is Cloudflare's observation on their own network, combined with their CEO's judgment, not a universally verified conclusion across the entire internet. This is an important topic, and the first section will focus on it specifically.
This issue affects a very fundamental thing: for the past few decades, websites have freely provided content, people came to watch ads, and advertisers paid for it. This cycle has sustained the entire internet. However, machines do not watch ads, so this payment mechanism is now becoming ineffective.
This article will explain three things separately: which are facts with direct sources, which are others' predictions, and which are my judgments. In the middle, it will stop to explain why different sources give different numbers for the same thing, because this topic is currently very hot and there are many misleading numbers. Finally, it will break it down into four layers and tell you what you can prepare for now.
I won't ask you to change anything tomorrow, but just to remind you that this transition has already started. The earlier you know, the more time you have to prepare.
- Those who are writing content, running websites, or building personal brands, and recently feel that traffic numbers are strange
- Those who help organizations create websites, knowledge bases, and annual reports
- Those who want to know how the evolution of 'AI marketing' will proceed
- A comparison table separating facts, others' predictions, and my judgments, with direct source links attached, which you can directly use to explain to others
- A habit of verifying trend numbers: first ask what is being measured, who is measuring it, and in which range it is measured
- A breakdown of four layers of opportunities (content layer, protocol layer, platform layer, service layer), with what you can do now for each layer
- A roughly ten-minute initial check to see if your content is now understandable to AI (including tool conditions and interpretation order)
According to Cloudflare, the amount of traffic generated by online robots has already exceeded that of humans. Their CEO originally estimated this would happen by the end of 2027, later revised it to early 2027, and finally said it occurred even earlier than that in 2026.
Over the past few decades, websites have provided content for free, people have viewed it while seeing ads, and advertisers have paid for it. This cycle has sustained the entire internet. Since robots do not view ads, this payment mechanism is now failing. Website costs are rising, while the sources of revenue are disappearing.
New money will come from three places: content licensing, exclusive details that can fill AI knowledge gaps, and a pay-per-use mechanism for content retrieval. The premise for all three things is the same: AI must find you, understand you, and be willing to cite you.
First, clearly explain the matter itself
The sources for this section are public statements by Matthew Prince, CEO of Cloudflare, and Cloudflare's official reports. Cloudflare is one of the world's largest network infrastructure companies, processing about 500 million requests per second across more than 350 cities and over 1,000 data centers, so its sample is substantial.
I will list the three things separately, and you should not mix them up.
- The crossover point has already occurred. Prince posted on June 3, 2026, saying: agentic traffic is growing too fast, and robot traffic has exceeded human traffic for the first time in internet history.
- Earlier than he himself predicted. In the same post, he wrote: originally thought it would be by the end of 2027, later revised it to early 2027, but it happened even earlier than that.
- The official version uses another term. The official report on July 1, 2026, wrote: 'more than 50% of internet traffic is already non-human traffic.' The scope of non-human traffic includes more than just AI agents.
- A hard number in terms of composition. The same report: as of June 2026, 52% of crawler requests were for AI training, whereas in spring 2025, this number was 22%.
- Paid agreements have existed for a long time, but were recently connected. The meaning of HTTP 402 is 'payment required.' Cloudflare's Pay Per Crawl uses it: AI crawlers request content, either by including a payment intent in the header to get a 200, or by receiving a 402 with a quote. It is currently in a closed beta test.
- The pricing range is expanding outward. Monetization Gateway aims to allow charging for websites, datasets, APIs, and MCP tools, initially using x402 for settlement in stablecoins. As of July 1, the status is open for pre-registration.
- x402 is not something belonging to Cloudflare. It is an independent open payment standard, entrusted to an independent foundation under the Linux Foundation, with AWS as one of the founding members.
Source link see Reference source at the end of the article.
The following few points come from a public interview's Chinese paraphrase and summary. I did not find a verbatim first-hand quote. He mentioned deductions and examples; these all lack a first-hand source for verification. Therefore, please treat them as his statements, rather than data references.
- 1000 times in five years. Robot traffic may reach 1000 times that of humans within five years. He explicitly stated this is a projection based on the current growth rate.
- A single request becomes thousands of visits. A person needs to look at about five websites to buy a camera, while an AI agent may visit 5000 websites to provide a better answer.
- The scale threshold is very high. If 1% to 10% of 500 million requests per second are suitable for micropayments, this system would need to support 10 million financial transactions per second on its first day. In comparison, Visa processes fewer than 100,000 transactions per second.
- A case of income reversal. He purchased the local newspaper in his hometown Park City, Utah, expecting that this newspaper's income from AI licensing this year would exceed display advertising.
- Organizations will become thinner. Management span will expand, and levels will decrease.
- Brand shortcuts will fail. For humans, brand is a shortcut for quick decision-making. AI agents have no patience limits and will continuously retrieve a large number of options, then judge based on price, location, and reviews.
- I estimate 100 times within two years, which is more conservative than 1000 times in five years, but once the breakthrough is made, it will grow quickly.
- I originally scheduled this matter for early 2027, but now I have moved the timeline forward to this year.
- The pace in Taiwan is roughly six to twelve months behind abroad, so taking action now is a very advanced position in Taiwan.
Different sources give different numbers for the same thing
This section I particularly want to write, because this topic is currently very hot, and there are many numbers circulating online, but a high proportion of them have unknown sources. When I was verifying, I encountered three things to be cautious about.
First, the four terms are not the same thing
Robot (bot), non-human traffic (non-human), AI crawler (AI crawler), and AI agent (AI agent) are four different classifications with different scopes. Cloudflare's official report said that non-human traffic has exceeded half, but it did not say that AI agents alone have exceeded humans. These two statements are very different.
Second, different companies' networks see different things
The 2026 report from the cybersecurity company HUMAN stated that in their observed AI driven traffic, AI agents and agent type browsers accounted for only 1.7%. The report explicitly stated this represents only the sample they saw on their platform. Fastly's 2026 report said that the growth rate of AI traffic in their network was 6.5 times that of human traffic. This growth mainly consisted of crawlers. These numbers are not contradictory; however, they cannot be used interchangeably.
Third, traffic share does not equal commercial value share
Cloudflare Radar measures the classification of HTTP requests. You cannot directly infer from 'robots make more requests' that 'human reading, staying, and consumption have been overtaken by machines.'
Why the old network business model will fail
How the old cycle works
This cycle has many problems, but it at least solves one very basic issue: who pays for content and data centers.
How to break this new situation
- Robots do not click on ads.
- AI takes website information and generates answers directly on its own interface, users even do not need to visit the original website.
- Websites bear the costs of writing, interviewing, data centers, and traffic, but do not receive traffic or ad revenue.
- At the same time, the frequency of network usage is increasing, and the cost of maintaining network operations is rising accordingly.
The result is a very contradictory situation: usage is increasing, but the payment mechanism is failing.
The criteria for value may change
The above four points are inferences about the mechanism. The following section is about 'how it may change', which has not been verified. I separate it from the above.
Prince uses Swiss cheese to metaphorically describe large language models: models already contain a lot of human knowledge, like the cheese itself, but there are many holes inside. He believes that the valuable content in the future will be the 'new knowledge' that fills these holes.
He gives the example of local newspapers. A person preparing to go to a ski town needs to know which hotel is better, which restaurant is authentic, and the local snow conditions. These small, specific, and only long-term accumulated local information are exactly what general models lack most.
This metaphor and example are both from secondhand paraphrasing, and I do not have the original first-hand sources. I place it here because it is easy to understand, not because it can prove anything. The following judgment is mine, and you may disagree.
Where is the opportunity: four layers
These four layers are stacked from bottom to top. For each layer, I first explain 'what is happening at this layer,' and then explain 'what you can prepare now.'
The things that general models can capture are losing value. Conversely, the details that have low traffic, aren't sensational, and are not captured by general models have become scarce under the new logic. Specifically, these include the actual practices of a particular industry, the real situation of a place, the cases you have served and their effectiveness numbers, the methodologies and judgment criteria you have accumulated over the years, and the lessons learned from the time you failed.
- Write out the details that only you have. Meeting minutes, activity records, post-class summaries, and questions customers have asked are the most effortless and continuous sources of content.
- But first, cross a boundary. Cases, performance numbers, meeting records, customer issues can only be written if the other party agrees to publicize them, you have done de-identification, and it does not involve confidentiality obligations. If unsure, ask the other party first, or only write the method without mentioning the subject. This step cannot be skipped.
- First have a trace, and do not worry about traffic. Without a trace, AI will not even have the opportunity to find you.
- If you are already doing short video content and have already started, then continue. If you have not started yet, I would suggest focusing your efforts on text content that can be crawled first, because that is the form AI can read.
Having content is not enough, AI needs to be able to read it, and it needs to read it efficiently. Cloudflare will convert complex HTML into more concise Markdown for websites that are willing to be crawled, allowing agents to read with lower costs, and also avoid irrelevant page elements filling up the model's context window.
Clean, readable content can reduce the agent's extraction cost, and there is a mechanism that can explain this. As for whether lower extraction cost will increase the likelihood of being cited, I have not seen any public data to confirm this. This is an assumption that needs to be tested yourself, not a confirmed sorting rule.
llms.txt: An entry guide for AI to enter the site. What you have, where the endpoints are, what the fields mean, and how to search.- Structured data (JSON-LD): Let machines understand that this is a class, a lecturer, an event, or a date.
- Main content uses clean, readable raw HTML, with an optional Markdown version available when necessary.
- Skill package: Write directly to the AI how to use your data to accomplish tasks.
Agent to Agent means both parties bring their own AI to accomplish tasks. Your AI talks to the other AI, compares prices, schedules time, filters potential partners, and then returns several options to you. For this to work, three things are needed: rules, trust, and ownership of computing power.
My own view is that the division of labor should be: humans handle trust and commitments, agents handle data and exploration, and the platform handles rules and order. Data is kept and updated by the individual, not concentrated on the platform. Computing power is paid for by the user, and the platform only provides algorithm mechanisms and trust mechanisms.
Think clearly about where your trust evidence is stored. Prince speculates that brand shortcuts will fail for AI because agents have no patience limits and will keep searching indefinitely. This is his speculation, and I have no data to confirm it.
Regardless of whether that speculation is accurate, one thing remains certain. Machines cannot read partners you have worked with, results you have produced, or how others describe you if these only exist in conversation records and memory. Turning trust, which originally exists only in relationships, into verifiable, accumulative, and citable records has value regardless of any prediction's proof.
For some organizations, the above three layers will be blocked by human resources, technical challenges, or methods of organization. They have content but no organization; they have websites but machines cannot read them smoothly; they know they need to prepare but don't know where to start. This gap often becomes something someone is willing to pay for.
If you are taking on clients, teaching, or doing consulting, this is a service line that will grow over time. The meaning of building a website is changing, and what is delivered is shifting from 'a nice-looking webpage' to 'a website that AI can find, read, and cite.'
Your preparation pace
This is my recommendation for individuals and small teams, not an industry timeline. The golden dot marks your current position. The sentence below each section is the completion criterion.
Seven things you can do first
| Actions | How to do it | How to know if it's done |
|---|---|---|
| Clarify what you're saying | One sentence to explain who you are, what problem you solve, and who it suits | This sentence should be something you can say out loud, and others can remember after hearing it |
| Focus on one main platform | Choose one platform to post on regularly, and use others as supplementary links | At least two posts per month, without interruption |
| Post meeting and event records directly | Organize into readable articles, without rewriting into long pieces | One post per event |
| Add entry points that AI can read | llms.txt, structured data, clean HTML | Accessible via browser, format validation passed |
| Buy a formal domain name | Use your own domain name, so the signals you accumulate stay under your name, not with the platform | Complete domain name setup and point it to your site |
| Do a citation test once a month | Ask several AI assistants with a fixed question, and record the question and results | You have a monthly record to compare |
| Before citing numbers externally, check the primary source | Find the original source link first | You need to provide the source, not just 'heard it' |
First confirm the tool conditions Use an AI that can open a URL or search the web. Different AIs vary a lot: some will actually open your URL, some only answer based on training data, and some use old cache. If the tool is wrong, the results are meaningless
Paste your URL and ask three questions
- Who is this person?
- What problems does he solve?
- In what situations should you consult him?
Remember four things each time Which tool, date, whether it actually opened your website (you can directly ask it 'Did you read this URL? Give me a quote from the original text'), and which page it referenced. Without these four pieces of information, you won't have anything to compare next time
How to interpret the results If the answer is accurate and specific, it means the AI understood this time. If the answer is vague and full of adjectives, it usually means the content itself hasn't been organized clearly. If it's wrong or can't answer, don't rush to blame your data structure. Check in order: ① Did it actually open your page? ② Is the page blocking scraping? ③ Is the search index indexed? ④ Is the content structure clear? ⑤ Are there contradictions between different pages?
A single result doesn't mean much. This is a check that needs to be run repeatedly and accumulate records to be useful. Three questions plus records, the first time takes about ten minutes. Testing with multiple AIs will take longer
Four things I don't recommend doing
- Don't do content farming-style mass production The new logic rewards content that fills knowledge gaps. Mass production here has no leverage
- Don't provide free AI Q&A services on your own website There are two reasons: the computing cost will be borne by you, and the other party won't upload their data, which means they also haven't accumulated their knowledge base. A more reasonable structure is to let the other party put your website into their own agent, and your data and their data exchange
- Do not use unverified numbers for external arguments. This topic is currently very popular, and many numbers circulate online; a high proportion of sources are unclear. I want to remind myself again: the column labeled 'he said' in this article, as well as the Swiss cheese and local newspaper examples, are all secondhand retellings that I did not trace back to primary sources. I keep them because they help with understanding. The argument in this article is supported by the few items with primary sources. The secondhand parts are for illustration; they are not used as evidence. When you write your own content, you can use this distinction: keep evidence and explanatory material separate.
- Do not treat licensing revenue as the premise of the plan. That matter is not in your hands.
The relationship between this article and my previous two articles
These three articles are three different perspectives of the same matter and can be cross-verified with each other.
In that article, I discussed: when AI starts to make decisions for people, marketing targets will include an AI. Marketing with AI treats AI as a tool, while marketing to AI treats AI as an audience. When I wrote that article, this was still a hypothesis, and my basis was the development direction of the A2A agreement.
Now there is external data. According to Cloudflare's view, the traffic from robots has already exceeded that from real people. This crossover point turns the hypothesis from 'this will happen in the future' into 'we can already see signs of it.' 'Marketing to AI' does not need to wait any longer.
At this current moment, you can learn 'marketing with AI' and also learn 'marketing to AI' →That article discussed the mechanism side: how an organization should inventory processes, establish a single source of truth, decompose tasks, clarify data boundaries and review points.
This article complements the motivation side: why we should do this now, and where the value will appear after it is done.
AI agents are about to take over work. Are your processes and knowledge prepared? →Conclusion
The special aspect of this turning point is that it arrived earlier than Cloudflare's original estimate, and it touched upon the payment logic at the very bottom of the network, not just an algorithm adjustment on a particular platform.
I don't think you need to panic about this matter, and I don't think there will be any immediate losses tomorrow. I just want to remind you: this matter has already started, and preparation takes time. Content trajectories need to accumulate, structures need to be organized, and citations need to be tested in practice. None of these can be completed within a week.