# Doxa Marketing Website # https://doxa.app # Content Signals (https://contentsignals.org/) — declare AI-use preferences # search: allow indexing for search engines # ai-input: allow AI assistants to cite this content in answers (GEO-critical) # ai-train: allow inclusion in AI training datasets Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: * Allow: / Allow: /api/media/ Disallow: /api/ # Course PDFs are not crawled directly, so search and AI funnel everyone to the # gated /courses/ pages, where details are captured before download. Disallow: /courses/*.pdf$ # /_next/static/ was Disallowed here until 2026-09-02. Removed on GSC evidence: # doxa.app serves /grace and /bible from the Next.js app, so blocking its JS and # CSS chunks stops Googlebot rendering those pages. The Pages report showed the # cost directly: 644 URLs "Blocked by robots.txt" and 12 "Indexed, though blocked # by robots.txt", every one a /_next/static/ chunk, plus a rendering handicap on # every Next-served page. Google's own guidance is to let crawlers fetch the # resources a page needs to render. Do not re-add it. # Preferred domain Host: https://doxa.app # Sitemap (unified index covering both Astro marketing pages and Next.js dynamic content) Sitemap: https://doxa.app/sitemap-index.xml # LLM context files # See https://llmstxt.org for specification # llms.txt = concise index, llms-full.txt = comprehensive product context Allow: /llms.txt Allow: /llms-full.txt # Block SEO scrapers (no value, just resource waste). # Standing rule (Garth 2026-07-18): never block a crawler that can get Doxa # found or cited in AI chat answers. YandexBot (Yandex Neuro/Alice answers) # PetalBot (Huawei Petal Search / assistant surfaces), and AhrefsBot (its # crawl powers the Yep.com search engine) were unblocked for that reason — # only scrapers serving NO user-facing search or AI answers stay here. User-agent: SemrushBot Disallow: / User-agent: MJ12bot Disallow: / User-agent: DotBot Disallow: / User-agent: BLEXBot Disallow: / User-agent: serpstatbot Disallow: / User-agent: DataForSeoBot Disallow: / # AI/LLM crawlers - welcome to all public content (GEO) # Note: every Disallow: /api/ block also explicitly Allow: /api/media/ so the # Postiz→Meta video proxy stays reachable. Per-UA blocks override User-agent: *, # so an Allow inside the wildcard group is not enough — each block needs its own. User-agent: GPTBot Allow: / Allow: /api/media/ Disallow: /api/ Disallow: /courses/*.pdf$ User-agent: Google-Extended Allow: / User-agent: CCBot Allow: / Allow: /api/media/ Disallow: /api/ Disallow: /courses/*.pdf$ User-agent: anthropic-ai Allow: / Allow: /api/media/ Disallow: /api/ Disallow: /courses/*.pdf$ User-agent: PerplexityBot Allow: / Allow: /api/media/ Disallow: /api/ Disallow: /courses/*.pdf$ User-agent: ClaudeBot Allow: / Allow: /api/media/ Disallow: /api/ Disallow: /courses/*.pdf$ User-agent: meta-externalagent Allow: / Allow: /api/media/ Disallow: /api/ Disallow: /courses/*.pdf$ User-agent: Bytespider Allow: / Allow: /api/media/ Disallow: /api/ Disallow: /courses/*.pdf$ User-agent: Applebot-Extended Allow: /