# robots.txt for credilinq.ai # Owner: Credilinq.ai # Purpose: Crawl directives for traditional search engines and AI agents. # Note for engineers: this file replaces the existing /robots.txt. Confirm any # current Disallow rules (admin paths, search, etc.) before deploying and merge # them in. WordPress typically auto-generates a robots.txt that already # disallows /wp-admin/ and similar; preserve those entries. # ----------------------------------------------------------------------------- # Default policy # ----------------------------------------------------------------------------- User-agent: * Allow: / Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php Disallow: /wp-includes/ Disallow: /wp-login.php Disallow: /xmlrpc.php Disallow: /feed/ Disallow: /comments/feed/ Disallow: /?s= Disallow: /?p= Disallow: /*?page_id= Disallow: /*?utm_source= Disallow: /*?utm_medium= Disallow: /*?utm_campaign= Disallow: /*?utm_content= Disallow: /author/ Disallow: /tag/ Disallow: /wp-content/plugins/ Disallow: /wp-content/cache/ Disallow: /wp-content/upgrade/ Disallow: /search/ Disallow: /thank-you # ----------------------------------------------------------------------------- # AI crawlers — explicitly allowed # Inference / search-time crawlers (low risk, high value) # ----------------------------------------------------------------------------- # OpenAI — ChatGPT User-agent: GPTBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: OAI-SearchBot Allow: / # Anthropic — Claude User-agent: ClaudeBot Allow: / User-agent: Claude-Web Allow: / User-agent: anthropic-ai Allow: / # Perplexity User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # ----------------------------------------------------------------------------- # AI training crawlers — review with Compliance before keeping these allowed. # Allowing means Credilinq content can be used to train future LLMs. # This is a deliberate business decision, not a default. # ----------------------------------------------------------------------------- # Google — Gemini training only (does NOT affect Google Search ranking) User-agent: Google-Extended Allow: / # Apple Intelligence training User-agent: Applebot-Extended Allow: / # Meta AI (Llama / Meta AI assistant) User-agent: meta-externalagent Allow: / User-agent: FacebookBot Allow: / # ByteDance — Doubao / TikTok AI User-agent: Bytespider Allow: / # Cohere User-agent: cohere-ai Allow: / # Common Crawl (training dataset used by many LLMs) User-agent: CCBot Allow: / # ----------------------------------------------------------------------------- # Sitemaps # ----------------------------------------------------------------------------- Sitemap: https://credilinq.ai/sitemap_index.xml # ----------------------------------------------------------------------------- # AI discovery file (not a robots.txt directive; referenced here for transparency) # ----------------------------------------------------------------------------- # llms.txt: https://credilinq.ai/llms.txt