Collection process
The collection process creates structured offering records from source information. Captured fields include the offering name, description, URL, course code, target audience, offering level, language, credential type, credit status and value, delivery format, schedule type, self-paced status, duration attributes, price attributes, skills and competencies, and topic category tags. Field availability depends on what the source provides and how that source is structured.
Normalization makes records
Normalization makes records usable across sources without assuming that source terminology is uniform. Delivery is represented as online, hybrid, or in person. Price information retains its amount, currency, cost type, payment pattern, and free status where those details are available. Duration can retain exact, minimum, maximum, and hour-based information. These distinctions matter when users compare programs with different presentation styles.
Credential forms classified
Credential forms are classified using the Credential Transparency Description Language vocabulary. The available classifications include Course, Certificate, MicroCredential, CertificateOfCompletion, LearningProgram, Diploma, Badge, ApprenticeshipCertificate, Degree, and CertificateOfParticipation. Classification provides a common vocabulary for filtering and market analysis while retaining the source record's own offering information.
Each offering receives
Each offering receives a dense vector embedding. Semantic search combines vector similarity with full-text search, allowing retrieval based on meaning as well as terms present in a source record. Users can further constrain results by province or country, source platform, credential form, and delivery mode.
Corpus distinguishes coverage
The corpus distinguishes between search coverage and dashboard coverage. Direct institutional crawl data is used for market-composition dashboards. Coursera and edX catalogue entries are searchable, but excluded from those statistics. Their subscription pricing models could distort price medians, so excluding them is a deliberate decision to preserve interpretability in market composition views.
Source limitations represented
Source limitations are represented rather than filled with assumptions. Platform catalogue entries provide a pricing model but not a dollar amount, and no dollar figure is displayed for those records. A substantial portion of platform catalogue records structurally lacks skills, duration, and price fields. This arises from scrape design and source structure, not from a failed collection run.
Absent information
CredScape does not treat absent information as a negative finding. A missing skill, duration, or price field indicates that the structured record lacks that source detail. Similarly, a collection timestamp identifies when the offering was collected, not a claim that the source page has remained unchanged since then.
Corpus inspectable
This methodology is intended to make the corpus inspectable. Decisions about source inclusion, field normalization, credential classification, retrieval, and dashboard scope are explicit so institutional readers can assess how the data fits their own planning and research questions.