AI Search Aug 8, 2026 15 min read

Google Lost the Scraping Case: What It Means for the Open Web—and for Your AI Search Visibility

A US judge rejected Google’s DMCA theory against a major SERP scraper. The deeper issue isn’t Google vs. SerpApi—it’s whether any business can rely on “anti-bot walls” to control AI crawlers. Here’s what changed, why it matters to SMEs and agencies, and how to build an intentional “open vs. closed” crawling policy that supports SEO, AEO, and GEO without betting your future on shaky legal ground.

Featured image for Google Lost the Scraping Case: What It Means for the Open Web—and for Your AI Search Visibility

Google losing a DMCA claim against a search-results scraper sounds like an inside-baseball SEO story. It isn’t. It’s a clean signal that the legal and technical tools most websites rely on to keep machines out may be weaker than we assumed—and that the “open web” principle cuts both ways.

I’m Marius Dosinescu, and from an AYSA.ai perspective I care less about who “won” this round and more about what this unlocks: more automated visitors, more AI agents, more scraping, more citation engines, and more businesses realizing they never explicitly decided what should be open, what should be closed, and what should be monetized differently.

This editorial is a practical guide to (1) what changed, (2) why it matters beyond Google, and (3) what SMEs and agencies should do now—especially if you want to stay visible in AI Search without giving away your entire business to bulk extraction.

Concise summary

Desk scene comparing copyrighted content protection vs. business model protection in a crawling policy discussion.
The core distinction: copyright protection vs. revenue protection.
  • A US judge tossed Google’s DMCA anti-circumvention claims against a scraper (SerpApi) because the anti-bot system Google referenced was framed as protecting ad revenue, not a copyrighted work.
  • The implications are broader than Google. Many sites deploy anti-bot layers primarily to protect revenue, operations, or “the business model.” That may not map cleanly to DMCA anti-circumvention protection.
  • AI agents and crawlers aren’t a niche. “Automated visitor access to public content” is foundational to the next web: AI Overviews, answer engines, shopping agents, research bots, and scrapers all sit on the same spectrum.
  • Businesses need an intentional crawler policy. Not “block everything,” not “leave everything open,” but a decision by page type, bot type, and business intent.
  • Execution is the moat. Strategy doesn’t help if the CDN, robots rules, or templates are misconfigured. AYSA’s model—monitor → prepare → request approval → execute—fits this moment.

Key takeaways (what to do with this)

Marketer explaining a simplified website stack that controls crawler access.
Your crawler policy already exists—it’s just accidental.
  1. Stop assuming your anti-bot stack equals legal protection. Treat it as an operational tool, not a legal shield.
  2. Define “open” vs. “closed” by intent. Discovery and citation need openness; bulk extraction may not.
  3. Segment your content. Your blog, product specs, and help docs may be “open”; pricing feeds, inventory endpoints, and high-value images may be “controlled.”
  4. Measure AI visibility like a business KPI. Monitor crawl patterns, citations, and the pages most likely to be summarized by AI.
  5. Move to approved, auditable execution. You want a record of what changed, why it changed, and who approved it.

Table of contents

Printed crawler access policy checklist held by a business owner.
A policy beats a panic-driven blocklist.

What happened: the Google vs. SerpApi moment

The trigger for this entire conversation was a court decision covered by Search Engine Journal: Google Lost Its Scraping Case – Now You Have To Pick A Side On The Open Web. The piece explains that Google’s DMCA claim against SerpApi (a company that scrapes and resells Google search results via API) was tossed—at least in the broad form Google attempted.

If you’re a business owner, you might ask: “Why should I care about a fight between two companies?” Because the logic used in a high-profile case becomes a template. Lawyers reuse arguments. Vendors market “DMCA-compliant blocking.” Site owners think they’re covered when they aren’t. And meanwhile, AI crawlers don’t wait for consensus.

This isn’t about cheering for a scraper or cheering for Google. It’s about the boundary conditions of the open web—and the operational reality that most websites are already negotiating that boundary with imperfect tools.

The ruling, in plain English: DMCA isn’t a business-model shield

Here’s the key idea from the SEJ coverage, translated for normal people:

  • The DMCA’s anti-circumvention rules are meant to protect copyrighted works and “technological measures” that control access to those works.
  • Google argued (as described in the SEJ article) that beating its anti-bot system (SearchGuard) should count as illegal circumvention.
  • The judge’s reasoning (again, per the SEJ coverage) was that the anti-bot wall described was protecting ad revenue / business operations, not acting like a lock around a copyrighted work. In other words: a wall around a business model isn’t the same as DRM on content.

That distinction matters because most anti-bot systems in the real world are deployed for:

  • rate limiting and server protection,
  • ad fraud prevention,
  • preventing competitive price scraping,
  • stopping inventory scraping and resale,
  • blocking content cloning for SEO or spam.

All legitimate business concerns. But not automatically “copyright locks” in the sense that DMCA anti-circumvention is built for. If you’ve been treating “we have a bot wall” as “we have a legal moat,” this is your wake-up call.

Why this matters to every site owner (not just SEO people)

For 20+ years, the internet’s default deal has been:

  • You publish content on public URLs.
  • Search engines crawl it.
  • You get discovery, traffic, and brand growth.

Now the same public URLs are being used for:

  • AI answer engines summarizing your pages,
  • AI Overviews referencing or paraphrasing your information,
  • agents that compare products and compile “best” lists,
  • scrapers feeding datasets, some legitimate, some not.

So the new deal businesses are trying to negotiate is:

  • “We want the upside of being open (discovery and citations)…”
  • “…without the downside (bulk extraction, republishing, and being used as unpaid infrastructure).”

This is exactly why the SEJ framing—“pick a side on the open web”—hits. You can’t keep pretending those two desires are automatically compatible.

The uncomfortable truth: you can’t praise the open web only when it benefits you

Most companies are inconsistent here, and I include marketers in that “most.” We say “open web” when we talk about:

  • ranking in Google,
  • earning links,
  • being cited by journalists,
  • getting included in “best tools” lists,
  • being recommended by AI assistants.

But we switch to “closed web” language when:

  • someone scrapes product pages to undercut prices,
  • an AI tool summarizes our blog and answers the query without a click,
  • a competitor clones our pages,
  • our images show up elsewhere.

Here’s my practical position: you don’t need a philosophical identity. You need a policy. Decide what you want the web to do for you—and what you refuse to subsidize. Then implement it technically, monitor it, and revise it as the ecosystem changes.

But you can’t outsource that decision to the courts, because the courts are reacting to ugly edge cases. And you can’t outsource it to plugins or CDN defaults, because those defaults were not designed for AI-era incentives.

From scraping to the agentic web: the same access question everywhere

One of the most important points in the SEJ piece is that the dispute is not merely a “scraping footnote.” It’s a preview of a broader shift: the agentic web.

In an agentic web, automated visitors don’t just read—they act:

  • a shopping agent checks availability, compares specs, and completes checkout,
  • a support agent searches your docs and drafts an answer,
  • a travel agent assembles packages based on your policies and rates,
  • a research agent extracts claims and builds a report.

Legally and operationally, many of these behaviors collapse into one category: automated access to public content. That’s why the “is it on the open web?” question matters more than the brand names involved in the lawsuit.

And it’s why businesses that want to win in AI search should stop thinking in terms of “SEO pages” vs. “AI pages.” It’s one web. The same URLs feed multiple discovery systems.

Risk map for SMEs: what can go wrong

SMEs feel this shift differently than big platforms. If you’re running a local clinic, a DTC brand, a small SaaS, or a regional services company, you don’t have an internal legal team or a bot-defense engineering group. You have a website, a marketing budget, and a need for predictable pipeline.

Risk 1: You block “bad bots” and accidentally block the future of discovery

Many anti-bot setups are blunt instruments. You might block whole cloud providers, whole geographies, or “unknown user agents.” But AI discovery can look “unknown” until it’s not. If you over-block, you may protect yourself from scraping today—and disappear from citations tomorrow.

Risk 2: You stay open and become free infrastructure

On the other side, staying fully open can mean:

  • increased crawl load,
  • content copied and outranking you,
  • pricing/specs scraped into competitor comparisons,
  • images reused without Attribution.

Risk 3: Your “anti-bot wall” becomes a conversion wall

Some businesses deploy defenses that hurt real users: CAPTCHAs on critical pages, aggressive JavaScript challenges, broken caching rules, or blocked accessibility tools. The business thinks it’s protecting itself; in practice, it’s taxing customers.

Risk 4: Compliance and reputation exposure

When you don’t control what gets extracted, you can end up with inaccurate summaries floating around the internet. Even if you didn’t authorize anything, your brand takes the hit. The fix isn’t “sue everyone.” The fix is “make your authoritative version easy to find, easy to cite, and hard to misinterpret.”

The real issue: you’re not deciding—your stack is

Most websites already have a crawler policy. It’s just accidental. It’s the emergent result of:

  • Robots.txt (often untouched for years),
  • CDN/WAF managed rules (often templated),
  • rate limiting (sometimes too strict, sometimes absent),
  • authentication and paywalls,
  • server caching behavior that unintentionally blocks or serves challenges,
  • CMS plugins that flip defaults without governance,
  • API endpoints that leak bulk data.

This is where the SEJ story becomes immediately practical: if courts are skeptical of framing “anti-bot” as “copyright protection,” then your primary lever is not legal—it’s operational and strategic.

And operational strategy needs ownership. Someone has to answer:

  • Which pages must be discoverable by search and answer engines?
  • Which pages can be summarized safely without harming revenue?
  • Which endpoints should never be crawled (pricing feeds, internal search results, session-driven pages)?
  • Which bots are allowed under which constraints (rate limits, attribution expectations, API alternatives)?

A practical crawler policy framework (open vs. closed, by intent)

Here’s the framework I recommend for SMEs and agencies. It’s simple enough to implement, but structured enough to survive the next 18 months of chaos.

Step 1: Classify your site by “value type”

Don’t start with “allow/deny.” Start with what the pages are:

  • Discovery content: blog posts, guides, category pages, service pages.
  • Decision content: pricing pages, comparisons, case studies, ROI calculators.
  • Operational content: account pages, checkout, internal search results, cart URLs.
  • Commodity specs: product specs, ingredients, dimensions, compatibility lists.
  • High-value assets: original photography, proprietary datasets, unique research.

Step 2: Decide what “open” means for each class

“Open” has levels:

  • Open for discovery (Indexing, basic Crawling)
  • Open for citation (AI can quote/summarize with attribution expectations)
  • Open for training/bulk extraction (often the controversial one)
  • Open via API (preferred if you want control, rate limits, and terms)

Many SMEs will land on something like:

  • Discovery + citation: mostly open
  • Bulk extraction: constrained
  • Operational: closed
  • High-value assets: controlled (watermarking, licensing, gated access, or selective openness)

Step 3: Implement with layered controls (not one lever)

Most businesses over-rely on one lever (robots.txt or WAF). You want layered controls:

  • Robots directives for good-faith crawlers and clear signals.
  • Rate limits to prevent high-frequency scraping from hammering origin servers.
  • Bot behavior rules: block abusive patterns (hundreds of requests per minute, aggressive parameter crawling), not just “unknown user agents.”
  • Content structure improvements so AI summarizes the right thing (clear definitions, FAQs, product attributes, citations).
  • API alternatives where bulk access is legitimate (affiliates, partners, marketplaces).

Important: this is not about “winning a war against bots.” It’s about making a deliberate trade: allow the forms of machine access that grow your business, and constrain the forms that extract value without a return.

Step 4: Make it auditable (governance)

If you can’t answer “what changed and who approved it,” you don’t have a policy—you have vibes. In a world where an aggressive block can destroy discovery, you need a controlled workflow.

This is where AYSA’s monitoring and approved execution philosophy becomes more than a product feature—it becomes operational hygiene.

Concrete SME scenario: an ecommerce catalog vs. AI scrapers

Let’s make this real with a scenario that mirrors what I see repeatedly.

Scenario

You run a mid-sized ecommerce brand selling specialty home goods. Your site has:

  • hundreds to thousands of product pages,
  • unique product photography,
  • help content (shipping, returns, care guides),
  • a small marketing team and a dev agency on retainer.

In the last year:

  • bot traffic increases,
  • your origin costs rise,
  • you notice your product descriptions appear on other sites,
  • at the same time, you want to be recommended in AI answers and shopping comparisons.

The common bad response

The panicked response is to “block everything AI-related,” or to deploy an aggressive WAF rule set that challenges any suspicious visitor. That can:

  • hurt Google crawling and indexation,
  • break shopping feeds,
  • reduce discoverability in answer engines,
  • create false positives for real customers (especially on mobile or corporate networks).

The better response: segment and control

A more resilient policy might look like:

  • Keep category pages and core product pages open for discovery (this is your shelf space).
  • Harden operational endpoints (cart, checkout, account pages) with strict bot controls.
  • Constrain abusive scraping patterns with rate limits and parameter handling rules.
  • Provide a clean, partner-friendly feed/API for legitimate bulk access.
  • Rewrite product content for “citable clarity”: structured specs, concise summaries, explicit policies—so AI systems cite the correct information.

This is exactly the kind of “SEO + technical controls + execution discipline” blend that determines whether AI search becomes a growth channel or a slow leak.

If you want a deeper approach to AI-era visibility, AYSA maintains practical resources at AI search visibility and AI SEO tools.

What agencies should rethink: deliverables, scope, and accountability

Agencies are caught in the middle of this shift. Clients want “rankings,” then “AI visibility,” then “stop scraping,” then “get cited,” then “block AI,” then “why did traffic drop?”

Here are the changes I believe agencies should make to stay credible:

1) Add “crawler policy” to the strategy layer

If your SEO strategy doesn’t include how automated visitors are treated, you’re missing a key variable. This isn’t just a technical add-on. It’s a business decision with brand and revenue impacts.

2) Shift reporting from only “rankings” to “presence”

As AI interfaces expand, visibility becomes multi-surface:

  • traditional blue links,
  • AI summaries,
  • citations in answer engines,
  • brand mentions across AI outputs.

Even if you can’t perfectly measure everything, you can monitor directional signals. Treat “AI mentions and citations” as part of the narrative, not a vanity metric. (SEJ has also been pushing this broader KPI shift in adjacent coverage and webinars; the ecosystem is moving.)

3) Sell execution with approvals, not “recommendations”

The age of PDF audits is ending. Clients don’t need 80 pages of “should.” They need controlled execution:

  • what to change,
  • why it matters,
  • risk level,
  • approval,
  • deployment,
  • monitoring and rollback.

That’s the thinking behind AYSA: we monitor, prepare changes, request approval, and execute accepted updates—so strategy doesn’t die in a backlog. If you’re evaluating models like this, start at pricing and the product approach described across our blog.

What to monitor now: signals that actually change decisions

SMEs often ask, “What’s the KPI for AI visibility?” The honest answer is: there isn’t one universal metric yet, and anyone promising a single number is oversimplifying.

But you can monitor a set of signals that lead to actions:

1) Crawl behavior changes (volume, paths, parameters)

Watch for sudden increases in requests to:

  • faceted navigation and filter parameters,
  • internal search URLs,
  • image directories,
  • API endpoints,
  • PDFs and downloads.

The action is not “block all bots.” The action is to refine rate limits, fix parameter handling, improve caching, and close endpoints that were never meant to be public.

2) Indexation and rendering anomalies

If your WAF starts challenging legitimate crawlers, you’ll often see indexation wobble. The action is to align technical controls with discovery goals—especially on your highest-value pages.

3) Citation readiness of your key pages

AI systems tend to summarize the clearest pages. You want your pages to be:

  • structured,
  • unambiguous,
  • consistent across templates,
  • explicit about definitions, specs, and policies.

The action is content + template improvements, not just “more blog posts.”

4) Brand presence drift

Even without perfect tooling, you can periodically check how your brand appears in AI-driven answers for your core topics. When misinformation appears, the fix is usually: update authoritative pages, improve internal linking, and ensure your best source is the easiest to cite.

AYSA’s approach is to make these monitoring inputs actionable through a controlled execution loop. Learn more at AYSA Monitoring.

How AYSA helps: monitor, prepare, ask for approval, execute

The open-web debate will continue. The lawsuits will continue. Meanwhile, your business needs operating procedures.

AYSA fits here in a straightforward way:

  • Monitor: detect changes in crawling patterns and visibility signals that indicate risk or opportunity.
  • Prepare: generate a prioritized set of site changes (technical, content, internal linking) aligned to AI search visibility.
  • Ask for approval: you stay in control—nothing ships without sign-off.
  • Execute: accepted changes get implemented, so strategy becomes reality.

This is especially important when the “right” move is not purely SEO or purely security—it’s both. Blocking, allowing, rate-limiting, restructuring pages for citations, clarifying policies, hardening endpoints: these are cross-functional decisions. Approved execution is how you keep them safe and auditable.

If you want to explore this model, start with:

What to do next (action list)

If you run a business site and want to stay visible in AI search without becoming a free data buffet, do this next:

  1. Write your crawler policy in one page. Classify pages: discovery, decision, operational, specs, assets. Define open vs. controlled for each.
  2. Audit your current reality. Check robots.txt, CDN/WAF rules, rate limits, and any “bot fight” plugins. Identify accidental blocks.
  3. Protect operational endpoints first. Lock down carts, account areas, internal search, and parameter-heavy URLs.
  4. Make your key pages more citable. Add clear summaries, structured sections, explicit policies, and consistent templates.
  5. Implement monitoring. You need alerts for crawl spikes and indexation anomalies—especially after security rule changes.
  6. Move to approved execution. Every change should have: reason, risk, owner, approval, and rollback plan.
  7. Revisit quarterly. The agentic web is evolving fast; your policy should too.

Sources and further reading

Note on sources: The supplied research context primarily includes the Search Engine Journal coverage above and site navigation links. Where this editorial discusses broader legal and operational implications, it is presented as analysis and practical guidance rather than claims of additional primary-source review.

Related AI SEO resources

Continue the AI search topic inside AYSA.

Use these pages to connect the article with AI SEO tools, AI visibility monitoring, AI Overviews and approved website execution.

Execution hubs

Turn this topic into a website action plan.

Use these AYSA hubs to move from reading to technical fixes, AI visibility monitoring, research, glossary context and approval-first SEO execution.

Marius Dosinescu, author at AYSA.ai

Written by

Marius Dosinescu

Marius Dosinescu is the founder of AYSA.ai, an entrepreneur focused on SEO automation, ecommerce growth, authority building and approved website execution for businesses that want organic growth without specialist overhead.

SEO execution, not more busywork

Turn SEO reading into approved website action.

AYSA monitors your website, prepares the work, asks for approval, and executes approved changes inside your website.

Start now View pricing

Only €29 to €99 per month, depending on the size of your business.

AYSA SEO Magazine

Latest search intelligence.

View all articles