CredScapeBot
Last updated 2026-09-27
CredScapeBot collects public information from institution websites about continuing education and professional development programs, including non-credit courses, certificates and micro-credentials. CredScape Market Intelligence Inc. uses this information to build market intelligence for continuing education teams. This page explains the requests you may see in your server logs, what we keep and how to stop crawling.
1. Identify CredScapeBot
Look for this User-Agent:
CredScapeBot/1.0 (+https://credscape.io/bot; bot@credscape.io)
Verify signed requests using HTTP Message Signatures (Web Bot Auth) and our public key directory.
2. What we collect and keep
We keep structured program facts and values we derive:
- Program title, institution and unit name.
- Price, currency, free status and pricing notes.
- Duration, delivery mode, schedule, dates and self-paced status.
- Credential type, course code, program type, offering format and course count.
- Province, country and whether a program is still offered.
- Topic classifications, derived skills tags and learning outcome counts.
- Source URL and collection date.
- Institution provider class, education type, host and offering counts.
We keep fetched HTML for at most 30 days after extraction, then delete it. Before any page text is sent to an AI model we remove site navigation, footers and other boilerplate with a content-extraction tool, and scrub email addresses and phone numbers. AI processing runs through OpenRouter with zero data retention, on providers that neither store prompts or outputs nor use them for training. Neither do we. Names in the body of a page may remain.
From wave 2 (October 2026):
We keep short excerpts of descriptions, recommended background and target audience text, with links to the source. We limit description excerpts to 200 characters and background and audience excerpts to 150 characters each. We also limit each excerpt to 30% of the original text, taking it from the beginning and ending at a word boundary.
We keep full copied prose, including learning objectives and skills lists, only during extraction and embedding processing, then delete it. We keep embeddings for internal use and do not return them to users.
We do not extract or retain instructor names.
We keep crawl logs for 24 months and opt-out records permanently.
3. How we crawl
We follow these rules:
- Obey robots.txt, checking the
CredScapeBotgroup before the*group, and honour publishedCrawl-delayinstructions. - Wait at least 3 seconds between requests to one host and make one request at a time per host.
- Refresh a site at most monthly.
- Stop crawling a host after a
401,403, bot challenge or CAPTCHA. Never retry with another identity, IP or tool. - Honour
Retry-Afteron429and503responses; otherwise, back off exponentially. - Treat an unreachable robots.txt file or a robots.txt server error as a full disallow until robots.txt responds.
- Do not use proxies.
We honour Content-Signal: ai-input=no by processing the page without sending it to a language model. We honour search=no by excluding the content from indexing and display.
4. Block CredScapeBot
Add these two lines to your site's robots.txt file:
User-agent: CredScapeBot
Disallow: /
5. Contact us or request removal
Email bot@credscape.io to ask us to stop crawling your site, remove text, correct a fact or remove personal data. Include the relevant website or page URL and describe your request.