What this online help covers:
This online help explains how to configure crawler access for search engines and AI assistants on your PWA. Learn about the 4 available presets in your back office, the comparison between traditional SEO and AI crawlers, the step-by-step configuration procedure, and the full list of supported crawlers.
Before you start:
Your robots.txt file is automatically generated with indexing directives tailored to your visibility goals and content protection preferences for AI models.
This online help explains how to configure crawler access for search engines and AI assistants on your PWA. Learn about the 4 available presets in your back office, the comparison between traditional SEO and AI crawlers, the step-by-step configuration procedure, and the full list of supported crawlers.
Before you start:
- Make sure your PWA has been published at least once under Publish > PWA > Publish .
- Check under Publish > PWA > SEO :: General parameters that the Authorize search engine to index the Progressive Web App option is enabled.
Your robots.txt file is automatically generated with indexing directives tailored to your visibility goals and content protection preferences for AI models.
1. Understand crawler access presets
GoodBarber lets you control precisely how search engines and AI assistants access your Progressive Web App (PWA). The system classifies crawler requests into 4 distinct categories:
What is the difference between "Maximum visibility" and "No AI training"?
- Real-time AI citations (AI answers): An AI assistant (like ChatGPT) opens your PWA live to answer a user, citing your site and sending traffic to your app.
- AI model training (AI training): Crawlers scrape your content to train future AI models, without generating direct visits to your site.
Which one to choose?
- Traditional SEO (Google, Bing): Users search using keywords. The Traditional SEO only preset keeps your Google ranking and indexing 100% active and unchanged.
- AI assistants (ChatGPT, Claude, etc.): Users ask questions to conversational AIs. The Traditional SEO only preset blocks AI assistants from reading or citing your app.
- Search engines: Traditional search engines (Google, Bing, etc.) that index your site to display it in search results.
- AI answers: AI assistants (ChatGPT, Claude, Perplexity, etc.) that cite or summarize your site's content in real-time answers.
- AI training: Crawlers that collect your content to build training datasets for future AI models.
- On-demand access: Temporary access when a user explicitly asks an AI assistant to open and analyze a specific URL on your site.
| Preset | Traditional Search Engines (Google, Bing) | AI Answers (ChatGPT, Claude, Perplexity etc.) | AI Model Training | Recommended Goal |
|---|---|---|---|---|
| Maximum visibility | ✅ Allowed | ✅ Allowed | ✅ Allowed | Maximize reach across all web search and AI platforms. |
| No AI training | ✅ Allowed | ✅ Allowed | 🚫 Blocked | Get cited by AI assistants and receive traffic without donating training data. |
| Traditional SEO only | ✅ Allowed | 🚫 Blocked | 🚫 Blocked | Preserve traditional Google/Bing SEO while requesting AI crawlers not to access your site. |
| Strict privacy | 🚫 Blocked | 🚫 Blocked | 🚫 Blocked | Request crawlers not to index or crawl the site (apps under development, testing, or internal use). |
What is the difference between "Maximum visibility" and "No AI training"?
- Real-time AI citations (AI answers): An AI assistant (like ChatGPT) opens your PWA live to answer a user, citing your site and sending traffic to your app.
- AI model training (AI training): Crawlers scrape your content to train future AI models, without generating direct visits to your site.
Which one to choose?
- Select Maximum visibility for full exposure across all search engines and AI platforms.
- Select No AI training to receive traffic from AI assistants while preventing them from training their models on your content.
- Traditional SEO (Google, Bing): Users search using keywords. The Traditional SEO only preset keeps your Google ranking and indexing 100% active and unchanged.
- AI assistants (ChatGPT, Claude, etc.): Users ask questions to conversational AIs. The Traditional SEO only preset blocks AI assistants from reading or citing your app.
2. Configure robots.txt in Guided mode
- Go to the menu Publish > PWA > SEO :: Robots.txtGuided mode, select the preset matching your needs under Choose a preset.
- (Optional) Click View details for this preset to check permissions (Allowed or Blocked) for each crawler category.
- Review the generated directives under GENERATED ROBOTS.TXT FILE PREVIEW.
- Click Save.
3. Configure robots.txt in Manual mode (advanced)
If your project requires custom crawling rules or specific directives, you can switch the editor to Manual mode:
- Switch the toggle at the top right to Manual.
- Edit the raw robots.txt file content directly in the text field.
- Click Save.
4. Full list of crawlers
Below is the detailed list of User-Agents identified and managed by GoodBarber presets, categorized by provider:
- Google: Googlebot, Google-Extended
- Microsoft: Bingbot
- DuckDuckGo: DuckDuckBot
- Yandex: YandexBot
- Apple: Applebot, Applebot-Extended
- OpenAI: GPTBot, OAI-SearchBot, ChatGPT-User
- Anthropic: ClaudeBot, Claude-SearchBot, Claude-User
- Perplexity: PerplexityBot, Perplexity-User
- Common Crawl: CCBot
- ByteDance: Bytespider
- Meta: Meta-ExternalAgent
- Amazon: Amazonbot
- Cohere: cohere-ai
- Diffbot: Diffbot
- Mistral: MistralAI-User