Promptwatch Logo

Funnelback

Funnelback crawls the websites and data repositories configured for a Squiz enterprise search collection.
Squiz - FunnelBackFunnelback
Search Engine Crawler

What is Funnelback?

Funnelback crawls the websites and data repositories configured for a Squiz enterprise search collection. It turns the retrieved documents into an index used by an organization's own website search or internal search. This is a scoped collection workflow, not a crawler building a general public web index.

Collection administrators decide which hosts and URL patterns belong in a crawl. Funnelback can follow links within that scope and, when the relevant option is enabled, read sitemap locations from robots.txt. A request in your logs usually means that a configured collection is discovering or refreshing material for its search results.

Squiz's crawler documentation says Funnelback honors the original robots.txt standard by default. Its support is narrower than that of some modern crawlers: Allow rules and path wildcards are not supported. An administrator can also configure a collection to ignore robots.txt after obtaining the site owner's permission. The bot metadata leaves compliance unconfirmed, and Cloudflare's bot directory marks it as not following robots.txt, so the effective behavior depends partly on the collection configuration.

No general AI search or model-training purpose is documented for this bot. A Funnelback index could be used by the organization that commissioned the crawl, but a visit does not make the page eligible for ChatGPT, Claude, or another unrelated answer engine. Blocking it mainly affects the organization's Funnelback search coverage.

Relevant for AI search

Is Funnelback relevant for AI search?

Yes. Funnelback collects pages for an AI product, so what it can crawl influences how AI systems describe your brand.

Funnelback feeds a search index that also powers that engine's AI answer features and overviews. One crawl can serve a classic results page and a generated summary, so blocking it costs you both the rankings you would expect and a growing share of AI answers built on the same index.

How to handle Funnelback

If your organization or a known partner uses Funnelback, allow the public paths that belong in that search collection and exclude account areas, search-result loops, and private repositories. A complete crawl block for the default agent is:

User-agent: Funnelback
Disallow: /

Squiz documents Funnelback as the default robots agent, matched case-insensitively. Its crawler normally honors Disallow, but a collection owner can change the agent or enable the documented ignore setting with the site owner's permission. Check access logs after a policy change, and use authentication or an edge rule when a restriction must be enforced.

Examples

  • A university lets Funnelback index its public course catalog while excluding the student account and internal search-result URLs.
  • An intranet administrator schedules a collection refresh, and the origin logs show Funnelback revisiting policy documents that changed since the previous crawl.
  • A partner's collection keeps requesting a retired section despite a new rule, so the site owner checks whether that collection was configured to ignore robots.txt.

Frequently asked questions about Funnelback

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

It retrieves the page for a configured enterprise search collection. The resulting index powers search for the organization that set up that collection.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard