scfc.Systems / field notes
Rebuilds and migrations / An information architecture essay

Schema starts before the JSON-LD

Structured data is the published expression of an entity model. Start by defining identities and relationships; the markup comes after the page has something coherent to describe.

A page can contain syntactically valid JSON-LD and still describe an ambiguous entity. Markup cannot decide whether two labels refer to the same service, which provider owns a course or whether a document applies to a product variant. Those decisions belong to the information model.

Name entities before properties

Identify the things the organisation needs to describe: products, services, providers, locations, people, documents or events. Define stable identifiers, preferred names, alternate labels, ownership and the relationships that matter. Be clear about distinctions. Two similar service names may represent separate eligibility rules; one product family may contain variants with different documents.

Then decide which source record is authoritative, how changes are governed and which pages present each entity. This connects editorial language, database records, URLs and published content.

Publish only what the page supports

Choose Schema.org types and properties that accurately describe visible, maintained content. Link identifiers consistently and include relationships only where the underlying records support them. Validate syntax, then compare the markup with the actual page and source data. A validator can catch formatting errors; it cannot confirm that a business relationship is true.

A practical inspection

  • Define entities and distinctions in business language.
  • Choose stable identifiers and authoritative source records.
  • Map relationships to pages and visible content.
  • Generate markup from maintained data where possible.
  • Validate syntax and separately verify the meaning.