# ========================================== # robots.txt for ishi-career.jp # # 注意: robots.txt は「最も一致する User-agent ブロック1つだけ」が # 適用される。個別ブロックを持つボットは User-agent: * を見ないため、 # 非公開パスの Disallow は各ブロックに明示する必要がある。 # ========================================== # ------------------------------------------ # デフォルト(個別指定のない全ボット) # ------------------------------------------ User-agent: * Disallow: /mypage/ Disallow: /admin/ Disallow: /offers/favorites/ Crawl-delay: 10 # ========================================== # 1. 検索エンジン(許可しつつ非公開パスは除外) # ※ Googlebot は Crawl-delay を無視する(Search Console で設定) # ========================================== User-agent: Googlebot Disallow: /mypage/ Disallow: /admin/ Disallow: /offers/favorites/ User-agent: Bingbot Disallow: /mypage/ Disallow: /admin/ Disallow: /offers/favorites/ Crawl-delay: 5 # ========================================== # 2. AI 検索・閲覧系クローラー(許可) # ユーザーの検索/質問を起点に来るため、流入につながる。 # 非公開パスは除外。純粋な学習収集系は Section 3 で拒否。 # ========================================== # OpenAI(ChatGPT 検索 / ユーザー起点アクセス) User-agent: OAI-SearchBot Disallow: /mypage/ Disallow: /admin/ Disallow: /offers/favorites/ Crawl-delay: 10 User-agent: ChatGPT-User Disallow: /mypage/ Disallow: /admin/ Disallow: /offers/favorites/ Crawl-delay: 10 # Anthropic(Claude のユーザー起点アクセス) User-agent: Claude-User Disallow: /mypage/ Disallow: /admin/ Disallow: /offers/favorites/ Crawl-delay: 10 # Perplexity(回答時に出典として引用 → 流入につながる) User-agent: PerplexityBot Disallow: /mypage/ Disallow: /admin/ Disallow: /offers/favorites/ Crawl-delay: 10 # Apple User-agent: Applebot Disallow: /mypage/ Disallow: /admin/ Disallow: /offers/favorites/ Crawl-delay: 10 # ========================================== # 3. 純粋な学習データ収集系クローラーの拒否 # 検索流入にはつながらず、モデル学習にのみ使われる。 # ========================================== # OpenAI(GPT 学習用) User-agent: GPTBot Disallow: / # Anthropic(Claude 学習用) User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / # Google(Gemini 学習用) User-agent: Google-Extended Disallow: / # Common Crawl(多くの AI が学習データとして利用) User-agent: CCBot Disallow: / # Meta (Facebook / Llama 学習) # meta-webindexer は 2026-09-03 の障害で見つかった Meta の新クローラ。 # Meta-ExternalAgent / FacebookBot とは別名を名乗るため、この2つだけでは素通りする。 # UA に developers.facebook.com のURLを含むが、SNS共有カード用の # facebookexternalhit とは別物なので、facebook でまとめて拒否しないこと User-agent: Meta-ExternalAgent Disallow: / User-agent: FacebookBot Disallow: / User-agent: meta-webindexer Disallow: / # Cohere / Omgili / Diffbot / ImageSift(その他 AI クローラー) User-agent: cohere-ai Disallow: / User-agent: Omgilibot Disallow: / User-agent: Omgili Disallow: / User-agent: Diffbot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: Timpibot Disallow: / # ========================================== # 4. SEO ツール・マーケティング系クローラーの拒否 # (負荷が高い代表的な SEOBot) # ========================================== User-agent: SemrushBot Disallow: / User-agent: AhrefsBot Disallow: / User-agent: MJ12bot Disallow: / User-agent: DotBot Disallow: / User-agent: rogerbot Disallow: / User-agent: BLEXBot Disallow: / User-agent: DataForSeoBot Disallow: / # ========================================== # 5. その他の大規模・攻撃的クローラーの拒否 # ========================================== # Amazon User-agent: Amazonbot Disallow: / # ByteDance(TikTok 等の運営元 / 攻撃的なクロールで有名) User-agent: Bytespider Disallow: / # Huawei(負荷が高いことが多い) User-agent: PetalBot Disallow: / # Yandex / Baidu(日本向けサイトでは流入が見込めず負荷のみ) User-agent: YandexBot Disallow: / User-agent: Baiduspider Disallow: / # ========================================== # Sitemap(全ボット共通 / ファイル全体で一度宣言すればよい) # ========================================== Sitemap: https://ishi-career.jp/sitemap.xml.gz