## **📌 Overview**

The Scholarly Scoring tool (found at **Left sidebar → Content → Scholar**) evaluates your content across multiple quality dimensions, including factuality, structure, and relevance. One of its core functions is determining what **type** of content it is analysing — such as a blog post, product page, or landing page — because each content type is scored against different benchmarks. This article explains how that classification works, why it sometimes produces unexpected results, and how factuality is evaluated.

## **🗂️ How Content Type Classification Works**

Scholarly Scoring uses a **weighted signal algorithm** that analyses three layers of your content to assign a content type. These layers are evaluated in order of descending weight:

1. **Metadata signals (highest weight — ~45%):** The tool first reads your page title tag, meta description, and URL slug. Keywords associated with commercial intent — such as "buy," "price," "model," "specs," or brand-plus-product-name patterns — push the classification strongly toward *product page*. Informational phrases like "how to," "guide," "what is," or "vs." lean toward *blog post* or *editorial*.
2. **Structural/heading signals (medium weight — ~35%):** The H1, H2, and H3 hierarchy is scanned for structural patterns. Product pages typically contain spec-list headings, feature-comparison subheadings, and short, noun-heavy section titles. Blog posts tend to use question-based or narrative subheadings. The density and formatting of heading text contribute significantly to the final classification.
3. **Body content signals (lower weight — ~20%):** Finally, the body text is analysed for trigger phrases, sentence structure, and paragraph length. Short paragraphs, bullet-heavy formatting, and imperative call-to-action language all reinforce a product-page classification.

The algorithm combines these three scores into a confidence percentage. If the confidence for any single content type exceeds **70%**, that classification is applied. Below that threshold, Scholarly Scoring falls back to the closest match and may flag the result with a low-confidence indicator.

## **🛒 Product-Page Signal Triggers to Watch For**

The following elements are the most common reasons a blog post gets misclassified as a product page. If your article contains several of these, the algorithm may tip past the 70% confidence threshold for *product page*:

- **Call-to-action (CTA) phrases** in body text or headings: "Buy now," "Get a quote," "Add to cart," "Request a demo," "Check availability."
- **Pricing language:** Dollar signs, percentage discounts, "starting at," "per month/year," or tabular price comparisons.
- **Specification-style formatting:** Tables or bullet lists with label–value pairs (e.g., "Weight: 2 kg," "Capacity: 500 units").
- **Schema-like heading patterns:** H2s that mirror product schema fields such as "Features," "Specifications," "Compatibility," "Reviews," or "Warranty."
- **Affiliate or commercial URL slugs:** Slugs containing "best," "top," "review," or "deal" combined with a product name push metadata signals toward commercial classification.

## **📝 Why Blog Posts Are Sometimes Misclassified**

Misclassification most commonly affects three blog-post formats:

- **Product roundups and "best of" lists:** These naturally contain spec comparisons, pricing references, and affiliate CTAs — the same signals the algorithm uses to identify product pages.
- **Tutorial content with tool recommendations:** In-depth how-to articles that feature embedded product CTAs or sponsored sections accumulate enough commercial signals to cross the classification threshold.
- **Venue-specific or location SEO content:** Pages targeting a specific venue, event space, or local business often include pricing tiers, capacity specs, and booking CTAs — structural patterns the algorithm currently maps to product or landing-page templates rather than local editorial content.

This is a **known limitation** of the current weighted model. The algorithm prioritises precision for e-commerce content but does not yet have a dedicated classification category for hybrid or local-SEO content types.

## **✅ How Factuality Scores Are Calculated**

Factuality scoring is independent of content-type classification and runs as a parallel evaluation. It measures how well your content is supported by verifiable, authoritative sources. The score is built from three components:

1. **External citation presence:** The tool checks whether your content links to or references sources outside your own domain. Content with zero external citations starts with a significantly reduced factuality baseline. For venue-specific and local SEO content, this means linking to official venue websites, local authority sources, or verified directories improves your score directly.
2. **Source authority signals:** When external links are detected, the algorithm evaluates the linked domains against authority benchmarks. Government, academic, and established news or industry domains contribute more to factuality than low-authority or unrelated domains.
3. **Claim-to-evidence ratio:** The body text is scanned for assertive claim language — statistics, superlatives, historical statements, or medical/legal/financial assertions. Each detected claim that is not followed by a citation or qualifying phrase reduces the factuality score incrementally.

For venue-specific SEO content in particular, factuality scores can drop if the page makes unverified claims about capacity, awards, or historical facts without linking to a primary source. Adding a simple reference link to the official venue website or a credible local directory is usually enough to recover several points.

## **🔧 How to Reduce Misclassification Risk Today**

While the algorithm is being improved (see below), you can reduce misclassification by adjusting the following before running a Scholarly Score:

- Rewrite your meta title and description to lead with informational intent language if your article is editorial, not commercial.
- Replace or soften CTA language in headings — move CTAs to a clearly labelled section near the bottom of the page.
- Avoid spec-style tables in the middle of blog content; use prose comparisons instead, or label the section explicitly as editorial context.
- For venue or local SEO pages, add a clear informational H1 (e.g., "A Complete Guide to [Venue Name]") rather than a conversion-oriented headline.

## **🗺️ Planned Algorithmic Improvements**

The Search Atlas product team has identified content-type misclassification as a priority improvement area. The following changes are on the roadmap:

- **Hybrid content-type category:** A new classification tier for roundups, local SEO pages, and affiliate editorial content that blends informational and commercial signals — reducing false product-page assignments.
- **Manual content-type override:** A setting inside Scholar that lets you specify your intended content type before scoring, ensuring benchmarks are applied correctly regardless of signal mix.
- **Venue and local SEO factuality templates:** Dedicated factuality rubrics that account for the structural conventions of local and hospitality SEO content, where pricing and specification data is editorial rather than transactional.

These improvements do not yet have a confirmed public release date, but they are actively in development. Updates will be announced in the Search Atlas product changelog.

## **💬 Need More Help?**

If you need further assistance, open the chat widget in the bottom-right corner of the platform and type **human teammate** to be connected with a member of our team.