DTC & B2B business research tool
DTC & B2B WebsitePublic Data Scanner
Give business owners and managers a faster way to review public DTC and B2B website data, including product count, sampled average price, blog volume, latest content update, publishing cadence, business model and market signals. Technology details remain available as supporting context. Website structures and public crawl access vary, so results are for reference only and cannot be guaranteed 100% accurate.
Methodology
Website Scanner Metrics and Detection Rules
This section explains each metric, calculation method and detection rule. Results come from public pages, links, sitemaps, source code and crawlable structured data. The scanner does not access private systems or grade overall website quality.
Website technology stack
Detects CMS platforms, ecommerce systems, headless CMSs and publicly verifiable frontend or backend technologies such as Shopify, WordPress, Storyblok, Next.js and Magento.
Uses official resource domains, platform-specific paths, DOM attributes and public variables. Confirmed layers may be combined; when no standard CMS is detected, the tool shows only the technologies it can verify and does not label the site custom-built.
Language versions
Lists publicly verifiable language versions. The scanner automatically selects the language branch used for the scan.
It scans the default page when it is English, otherwise prefers a verified English version on multilingual sites. If no English version can be verified, it scans the site's default language.
Analytics tools
Detects public installations of tools such as Google Analytics, Google Tag Manager and Microsoft Clarity.
Requires explicit script URLs, frontend variables or tracking IDs. A Google Search Console verification tag proves ownership only and is not treated as an active analytics runtime.
Advertising and conversion
Detects Google Ads, Floodlight, Meta Pixel, TikTok Pixel and related advertising infrastructure.
Pixels require explicit code signals. Domain verification and Shopify Web Pixels are reported separately and do not prove a specific advertising pixel is active.
Plugins and components
Detects public third-party components such as Klaviyo, Yotpo, Judge.me, GoAffPro, Gorgias, Intercom and Zendesk.
Only identifiable vendors are listed. A generic live-chat label is treated as a contact method rather than a named plugin.
Affiliate links
Checks for public affiliate, referral, partner or ambassador program entry points.
A real public link or provider endpoint is required. A word mentioned in body copy alone does not count as an affiliate program.
Contact methods
Detects public email, phone, WhatsApp, live chat, contact forms and feedback paths.
The scanner samples the homepage, selected product pages and contact-related pages. It reports method types rather than exposing full personal contact details.
Business model
Estimates whether the site is closer to DTC, B2B, DTC + B2B or another model.
DTC signals include products, prices, carts and checkout. B2B signals include quotations, wholesale, distributors, downloads and industry solutions. Affiliate links do not count as B2B evidence.
Public URLs
Counts URLs discovered from sitemaps, homepage links and public navigation.
This is a discovery estimate, not the site's database total or indexed-page count. Partial coverage is disclosed when a scan boundary is reached.
Product pages
Estimates the number of publicly discoverable product pages.
Uses URL paths, sitemap context, structured data and common ecommerce patterns such as /products/, /product/ and Product JSON-LD.
Category pages
Estimates product collections, categories and catalog pages.
Uses paths such as /collections/, /category/ and /catalog/. Brands often mix collections, campaigns and categories, so the result is approximate.
Sampled average price
Extracts prices from a product sample and calculates an average. It is not the true average of every product on the site.
For Shopify, the scanner prioritizes the first 10 non-gift-card products in the public Best selling order and uses the merchant's public base market and currency. It falls back to distributed sitemap sampling when that order is unavailable. Three to seven valid samples use a direct average; eight or more remove one high and one low value.
Blog and news
Estimates public blog, news and resource pages.
Uses paths such as /blog/, /news/, /articles/ and /resources/ together with sitemap context. It measures discoverable content assets, not content quality.
Latest content
Shows the latest verified publication date across sampled blog, news and resource content.
Publication dates come from article pages or same-site RSS/Atom feeds. dateModified and sitemap lastmod are used only as clearly labelled update-date fallbacks.
Publishing cadence
Estimates recent publishing frequency across blog, news and resource content.
Requires at least three distinct verified publication dates and uses the median interval. Modified dates, listing-card dates and sitemap lastmod do not enter the cadence calculation.
Authors
Counts distinct authors found in sampled public content.
Uses author meta, Article JSON-LD and explicit bylines. A result of zero means no author was exposed in the sample, not that the site has no content team.
Social links
Detects public LinkedIn, YouTube, Instagram, TikTok, Facebook, X and Pinterest links.
The result records public entry points only; it does not assess account activity or audience size.
Evidence
Provides representative public URLs or source signals for scanner findings where possible.
Only a small evidence sample is shown. Dynamic scripts may mean evidence appears as a URL, script source, page field or sampled page.