# ───────────────────────────────────────────────────────────────────────────── # robots.txt # ───────────────────────────────────────────────────────────────────────────── # Nothing is blocked. The brief's whole thesis is that this content should be # crawled, understood and cited. A site that exists to be a knowledge layer # has no reason to disallow retrieval. # # Search and internal utility endpoints are allowed too — /search/ is a real # page with a real URL and its results are legitimate landing pages. # # NOTE ON llms.txt (§37 of the brief): no llms.txt is published. The brief is # explicit that it is not a substitute for crawlability, structure, evidence # or authority. Crawl budget here is trivial, so no crawl-delay is set. # ───────────────────────────────────────────────────────────────────────────── User-agent: * Allow: / # AI retrieval and answer systems — explicitly welcomed, not merely tolerated. User-agent: GPTBot Allow: / User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: PerplexityBot Allow: / User-agent: ClaudeBot Allow: / User-agent: Claude-Web Allow: / User-agent: Google-Extended Allow: / User-agent: Applebot Allow: / User-agent: Applebot-Extended Allow: / User-agent: Bingbot Allow: / # Sitemap index (§23). Children: /sitemaps/pages.xml, articles.xml, fabrics.xml, # countries.xml, glossary.xml, research.xml, datasets.xml, manufacturers.xml Sitemap: https://kubramal.com/sitemap.xml