# PROPOSED physical robots.txt for covid19resources.ca # Place at /var/www/html/robots.txt (a physical file overrides WordPress's virtual robots.txt). # Preserves the current virtual rules, adds Events Calendar crawl-trap disallows. User-agent: * Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php # --- The Events Calendar crawl-trap: stop bots walking the infinite date/filter URL space --- Disallow: /*?*tribe-bar-date= Disallow: /*?*eventDate= Disallow: /*?*eventDisplay=past Disallow: /*?*eventDisplay=list Disallow: /*?*tribe_eventcategory= Disallow: /*?*ical=1 Disallow: /*?*outlook-ical= Disallow: /*?*post_type=tribe_events&* # FilterBar param space (plugin is active — expands the trap further): Disallow: /*?*tribe_events_cat= Disallow: /*?*tribe-bar-* # WPML language-duplicated calendar paths (translated slugs multiply the trap ~3x): Disallow: /*?lang=*&*tribe-bar-date= Disallow: /events/*/past/ Disallow: /evenements/*/past/ Sitemap: https://covid19resources.ca/sitemap.xml Sitemap: https://covid19resources.ca/sitemap.html # NOTE: PetalBot / serpstatbot / Scrapy frequently ignore robots.txt. # robots.txt only helps the well-behaved crawlers (Semrush, DotBot, SERanking). # The Cloudflare cache rule (see proposed-cloudflare-cache-rule.md) is the real fix.