How we designed our study of the Belgian SME website

The State of the SME Website in Belgium 2026 measures the technical, commercial and generative (GEO) readiness of discoverable Belgian SME websites. This article explains the design: what we measure, how we find companies and why we publish our own weaknesses.

The State of the SME Website in Belgium 2026 is a national baseline study of Belgian SME websites by SEMANU. It is a census, not a sample: every discoverable SME website is measured with a deterministic, public instrument. There is no AI in the measurement. The code is open source, the pseudonymised dataset is released under CC BY 4.0 and per-company results are never published, only aggregates.

In short

  • A census of every discoverable Belgian SME website, built on open data from the Crossroads Bank for Enterprises.
  • Three discovery layers: the register's website field, the Common Crawl web archive and domains derived from the company name, accepted only with proof of ownership.
  • The instrument is deterministic and public: anyone can run the same code on the same site and get the same result.
  • The selection effect of web research is measured and published per discovery layer, not argued away.
  • Protocol, codebook and dataset are published with a permanent DOI; nothing is published per company.

Why this study exists

Anyone asking how Belgian SMEs are doing online today mostly finds sales talk: figures from parties selling websites, tools or advice, measured in ways nobody can verify. We wanted a baseline you can recompute instead of believe. So we publish the full instrumentation alongside the results: the protocol, the measurement code and the thresholds behind every indicator.

A census, not a sample

The study does not measure a thousand random sites but every SME website it can find. That shifts the core question from sampling error to coverage: not “how representative is the sample” but “what share of the population do we see”. We report that coverage explicitly, with an independent recall check against a list that played no part in finding the sites.

How we find companies with a website

The backbone is the enterprise register. On top of it, three layers look for websites: the website field companies declared to the register, the public Common Crawl web archive in which we recognise enterprise numbers on pages and thirdly domain names derived from the registered company name: a candidate only counts once ownership is proven. A fourth layer through a search API died along the way: Google closed that route early in 2026. Rather than hiding the gap, we replaced it with a validation step that measures how much we miss.

What we measure and why no AI is involved

Per website the scanner records only machine-verifiable indicators: reachability, security, findability, contact options, commercial signals and static GEO readiness: can an AI crawler read and understand the site? No judgement, no model, no interpretation: the same code on the same site always gives the same answer. The instrument is versioned and every measured row carries the version that produced it. We discarded version 1.0 entirely after a test on our own website exposed five detectors that looked reasonable but were almost always right or almost always wrong. Those defects are described in the version history, with tests that demonstrate the old behaviour.

The selection effect: named, bounded, published

We raise the sharpest objection to this design ourselves: measuring discoverable websites means measuring companies selected for discoverability. Our answer is not denial but bounding. All discovery layers select in the same direction, towards stronger web presence, so every shortfall we report is a lower bound on the true shortfall. Every core figure is broken down by discovery layer, the overlap between layers is measured and a random sample of undiscovered companies is checked by hand. The design's biggest unknown becomes a measured figure with an interval, not a footnote.

Everything verifiable

The measurement code is open source under the MIT licence. The protocol is preregistered and archived with a permanent DOI and the pseudonymised dataset (all indicators per enterprise under a random study ID, without name, VAT number or URL) is released under CC BY 4.0. The scanner identifies itself honestly with its own name and contact address, honours robots.txt, fetches at most four pages per site and spaces its requests. The ethical framework follows the Menlo Report and the ALLEA European Code of Conduct for Research Integrity.

What comes next

The scan runs in waves over several weeks. First results follow here: how Belgian SMEs are really doing technically, commercially and in the AI search era, with every limitation stated. If you do not believe it, you can recompute it. As it should be.

Questions about this study, or interested in collaborating? SEMANU Research welcomes methodological feedback, media partnerships and sector organisations that want to use the results.