# robots.txt for www.smaply.com # Updated 2026-09-01. Goal: cut bulk crawler bandwidth WITHOUT costing acquisition. # # THE RULE THIS FILE FOLLOWS. Two questions, in this order: # 1. Can this crawler send us a human, or is it fetching on a live human's # behalf? If yes, it is an acquisition channel. ALLOW IT. No exceptions for # bandwidth. Search engines, AI answer engines and user-triggered fetchers # all fall here. # 2. Otherwise, is it building a training corpus or reselling scraped content? # Then it can never send us anyone, and blocking it costs nothing. BLOCK IT. # Bandwidth savings come entirely from question 2. The heavy crawlers (GPTBot, # CCBot, ClaudeBot, Bytespider) are all in that group, and the referral-driving # crawlers are low-volume by design, so there is no trade-off to make here. # If you are ever unsure which side a new crawler is on, ALLOW it and measure. # # HOW THIS FILE IS READ: a crawler obeys ONLY the single most specific User-agent # group matching its token. Groups are not inherited. Anything not named below # falls into the "*" group and is fully allowed. # # ORDERING MATTERS: Applebot-Extended and YandexAdditional are blocked while # Applebot and Yandex stay allowed. Because the allowed token is a prefix of the # blocked one, those narrow groups are deliberately placed FIRST. Google, Bing and # Apple resolve by longest match and would be fine either way, but simpler parsers # take the first match. Add any new "-Extended" style token above the allow groups. User-agent: * Disallow: # Training-only opt-out tokens for vendors whose search crawlers we welcome. # Neither is a crawler that can cite us: Google confirms Google-Extended is not # used for Search inclusion or ranking and does not affect AI Overviews or AI # Mode (those run on Googlebot), and Applebot-Extended does not crawl at all. User-agent: Google-Extended User-agent: Applebot-Extended Disallow: / # Yandex's separate AI-training token. The Yandex search crawler below is # unaffected and keeps indexing us normally. User-agent: YandexAdditional User-agent: YandexAdditionalBot Disallow: / # Search engines and Google's own tooling: full access, always. Organic search is # our main acquisition channel. Listed explicitly as a guardrail, since AdsBot-Google # ignores the "*" group and an explicit group survives a careless edit to "*". User-agent: Googlebot User-agent: Googlebot-Image User-agent: Googlebot-News User-agent: Googlebot-Video User-agent: Google-InspectionTool User-agent: Storebot-Google User-agent: AdsBot-Google User-agent: GoogleOther User-agent: Bingbot User-agent: msnbot User-agent: DuckDuckBot User-agent: Applebot User-agent: Yandex User-agent: Baiduspider User-agent: PetalBot User-agent: Bravebot Allow: / Disallow: # AI answer engines: these build the indexes assistants cite at answer time, so # they are the reason we appear in ChatGPT, Claude and Perplexity answers at all. # ChatGPT alone sends ~80% of our AI referral traffic. Blocking these would cost # referrals and save almost no bandwidth, because they crawl far less than the # training bots. Do not move any of these into a block group. User-agent: OAI-SearchBot User-agent: Claude-SearchBot User-agent: PerplexityBot User-agent: DuckAssistBot User-agent: YouBot User-agent: PhindBot User-agent: iAskBot User-agent: iaskspider User-agent: LinerBot User-agent: Andibot User-agent: NotebookLM User-agent: Gemini-Deep-Research User-agent: Google-CloudVertexBot Allow: / Disallow: # User-triggered fetchers: these fire only when a real person pastes our URL into # an assistant and asks about it, or points an agent at us. Highest-intent traffic # we get, and negligible bandwidth since there is one fetch per human request. # Blocking these makes us the site that "cannot be read" to an evaluating prospect. User-agent: ChatGPT-User User-agent: Claude-User User-agent: Perplexity-User User-agent: MistralAI-User User-agent: Kimi-User User-agent: Meta-ExternalFetcher User-agent: Operator User-agent: NovaAct Allow: / Disallow: # Link-preview / unfurl bots: full access, or shared links lose their preview card # on Facebook, Instagram, WhatsApp, X, LinkedIn, Slack, Telegram and Discord. User-agent: facebookexternalhit User-agent: Twitterbot User-agent: LinkedInBot User-agent: Slackbot User-agent: Slackbot-LinkExpanding User-agent: WhatsApp User-agent: TelegramBot User-agent: Discordbot Allow: / Disallow: # --------------------------------------------------------------------------- # From here down: bulk training corpora and commercial scrapers. None of these # can cite us or send us a visitor, so blocking them costs no acquisition. This # is where the entire bandwidth saving comes from. # --------------------------------------------------------------------------- # OpenAI and Anthropic model-training crawlers. Independent of their search and # user tokens above, so blocking these does not affect ChatGPT or Claude citations. # anthropic-ai and Claude-Web are retired tokens, kept for older crawler builds. User-agent: GPTBot User-agent: ClaudeBot User-agent: anthropic-ai User-agent: Claude-Web Disallow: / # Open training-corpus builders. CCBot (Common Crawl) is the single largest # upstream source for third-party model training, so one block here travels far. User-agent: CCBot User-agent: AI2Bot User-agent: Ai2Bot-Dolma User-agent: AI2Bot-DeepResearchEval User-agent: LAIONDownloader User-agent: laion-huggingface-processor Disallow: / # Bulk image and media harvesters. Images are ~69% of our served bandwidth, so # this is the group aimed most directly at the overage. User-agent: ImagesiftBot User-agent: imageSpider User-agent: img2dataset Disallow: / # ByteDance: among the heaviest crawlers on the web and a known robots.txt # violator, so expect this to need an edge rule to actually be enforced. User-agent: Bytespider User-agent: TikTokSpider Disallow: / # Meta's training and indexing crawlers (Meta-ExternalFetcher, the user-triggered # one, is allowed above). Honours robots.txt unreliably, so best-effort only. User-agent: Meta-ExternalAgent User-agent: meta-webindexer User-agent: FacebookBot Disallow: / # Amazon's ingestion crawlers for Alexa, Rufus and Bedrock. Zero measured referral # traffic to us in the last 12 months, so the crawl budget buys us nothing. User-agent: Amazonbot User-agent: Amzn-SearchBot User-agent: Amzn-User User-agent: bedrockbot User-agent: amazon-kendra User-agent: amazon-QBusiness Disallow: / # Other vendors' model-training crawlers. Their user-facing fetchers # (MistralAI-User, Kimi-User) are allowed above; these bulk trainers are not. User-agent: cohere-ai User-agent: cohere-training-data-crawler User-agent: DeepSeekBot User-agent: PanguBot User-agent: TongyiBot User-agent: ChatGLM-Spider User-agent: YiyanBot User-agent: WRTNBot User-agent: SBIntuitionsBot Disallow: / # Crawl-for-hire and scraping infrastructure. These resell our content through an # API and answer to no search product, so there is no citation to lose. User-agent: Diffbot User-agent: Scrapy User-agent: Timpibot User-agent: FirecrawlAgent User-agent: Crawl4AI User-agent: Crawlspace User-agent: ApifyBot User-agent: ApifyWebsiteContentCrawler User-agent: ExaBot User-agent: ExaSearchBot User-agent: TavilyBot User-agent: VelenPublicWebCrawler User-agent: Panscient User-agent: Devin Disallow: / # Webz.io data resale. Omgilibot is being retired in favour of Webzio-Extended, # so all three tokens are listed to cover the transition. User-agent: omgili User-agent: omgilibot User-agent: Webzio-Extended Disallow: / # Semrush's AI-content crawlers only. Plain SemrushBot is left allowed via "*" # so our own site audits and position tracking keep working. User-agent: SemrushBot-OCOB User-agent: SemrushBot-SWA Disallow: / Sitemap: https://www.smaply.com/sitemap.xml Sitemap: https://www.smaply.com/sitemap.xml