# robots.txt — merci-michel.com # Crawl rules per RFC 9309 (https://www.rfc-editor.org/rfc/rfc9309). # Crawlers may access the whole site; only the internal generator tool and the raw # template partials are excluded. # # Content-Signal (https://contentsignals.org) declares how the listed crawlers may # USE this content: search = search indexing (allowed); ai-input = use as live AI # input such as answers/grounding (allowed); ai-train = AI model training (not # allowed). These are usage preferences, separate from the crawl rules above. Sitemap: https://www.merci-michel.com/sitemap.xml User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /generator Disallow: /shared/ # AI assistants, AI search, and training crawlers — explicitly welcomed, # with the same two exclusions as everyone else. # (Mirrors App::AI_CRAWLER_PATTERN in server/App.php — keep in sync when adding crawlers. # Not 1:1: policy-only tokens like Applebot-Extended have no live-UA match in that regex.) User-agent: GPTBot User-agent: ChatGPT-User User-agent: OAI-SearchBot User-agent: ClaudeBot User-agent: Claude-Web User-agent: Claude-User User-agent: anthropic-ai User-agent: Google-Extended User-agent: PerplexityBot User-agent: Perplexity-User User-agent: CCBot User-agent: Applebot-Extended User-agent: Amazonbot User-agent: Bytespider User-agent: Meta-ExternalAgent Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /generator Disallow: /shared/