Dataset schema: getting research data listed in Google Dataset Search

Dataset schema describes a body of data, such as a CSV of survey responses, a set of sensor readings or a folder of trained model files, so that Google Dataset Search can index it. Google requires just two properties, name and description. The recommended ones (license, creator, download links, coverage dates) are what make a listing worth clicking.

Know where it shows up before you invest. In November 2025 Google announced it was phasing dataset features out of ordinary search results and updated its documentation to say Dataset markup feeds Dataset Search only. If researchers, journalists or data teams look for your files, the markup still earns its place. If you hoped a blog post with a chart would get a rich result, it won't.

JSON-LD example

Copy this into a <script type="application/ld+json"> tag and replace the values with your own.

json
{
  "@context": "https://schema.org",
  "@type": "Dataset",
  "name": "Riverside Allotment Soil Tests, 2021-2025",
  "description": "Soil pH, organic matter and lead readings from 64 community garden plots, sampled each spring and autumn by volunteers and analyzed by a university extension lab.",
  "url": "https://example.com/data/allotment-soil-tests/",
  "identifier": "https://doi.org/10.1234/example.soil.2025",
  "keywords": [
    "soil chemistry",
    "urban agriculture",
    "lead contamination"
  ],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "isAccessibleForFree": true,
  "version": "3",
  "creator": {
    "@type": "Organization",
    "name": "Riverside Growers Association",
    "url": "https://example.com/"
  },
  "temporalCoverage": "2021-04-01/2025-10-31",
  "spatialCoverage": {
    "@type": "Place",
    "name": "Riverside Allotments, Dayton, Ohio"
  },
  "distribution": [
    {
      "@type": "DataDownload",
      "encodingFormat": "CSV",
      "contentUrl": "https://example.com/wp-content/uploads/2025/11/soil-tests-2021-2025.csv"
    }
  ]
}

Properties

PropertyStatusWhat it is
nameRequiredA distinct, descriptive title. Two different datasets should never share a vague name like "Snow depth".
descriptionRequiredA summary between 50 and 5000 characters. Markdown is allowed; in JSON-LD, write line breaks as \n.
creatorRecommendedThe Person or Organization that produced the data. Google suggests an ORCID iD in sameAs for people and a ROR ID for institutions.
licenseRecommendedA URL for one specific license version, such as a particular Creative Commons deed, not a generic licensing page.
distributionRecommendedOne DataDownload per file, each with contentUrl and an encodingFormat such as CSV or XML. contentUrl is required inside each DataDownload.
identifierRecommendedA DOI or Compact Identifier. List several as a JSON array when the data has more than one.
temporalCoverageRecommendedThe period covered, as an ISO 8601 date or interval. Two dots (..) mark an open end.
spatialCoverageRecommendedA place name, a point or a GeoShape box, included only when the data has a geographic dimension.
isAccessibleForFreeRecommendedtrue when anyone can download the data at no cost.
sameAs / isBasedOnRecommendedsameAs names the canonical copy when you republish someone else's data unchanged; isBasedOn marks data you changed significantly or merged from several sources.
includedInDataCatalogRecommendedThe DataCatalog, such as a repository, that the dataset belongs to.

What counts as a dataset

Google's definition is wide. A table or a CSV file qualifies, and so does an organized collection of tables, a proprietary file that holds data, a group of files that only make sense together, images that capture data, and machine learning files such as trained parameters.

The markup describes the data, not the data itself. It says what the dataset is about, who made it, which period and place it covers and where to download it. It never lists the actual values.

Put the markup on the dataset's landing page: the one page that exists to describe it. Repositories also have listing and search pages that mention many datasets. Those don't need markup, and if you add it there anyway, point sameAs at the landing page so Google knows which description is canonical.

Republished and derived data

Open data gets copied, cleaned and combined constantly, and Google asks you to say which case you are in:

  • A straight republication of someone else's dataset: add sameAs with the original's canonical URL. Each sameAs value must identify one dataset only.
  • A version you changed significantly, or one built from several sources: use isBasedOn for each original.
  • Persistent identifiers: attach DOIs or Compact Identifiers with identifier, repeated as a list if there are several.

Keep text properties under 5000 characters, because Dataset Search ignores anything past that point. Names and titles should be a few words, not a paragraph.

Publishing data from a WordPress site

A common setup is a page that explains the dataset plus the files uploaded to the Media Library. The page is the landing page; each attachment URL becomes a contentUrl in distribution.

Hydrogen SEO's schema screens don't offer a Dataset type, so print it with the hydrogen_seo_schemas filter, which receives every schema object before output:

php
add_filter( 'hydrogen_seo_schemas', function ( $schemas ) {
    if ( is_page( 'allotment-soil-tests' ) ) {
        $schemas[] = [
            '@context'    => 'https://schema.org',
            '@type'       => 'Dataset',
            'name'        => 'Riverside Allotment Soil Tests, 2021-2025',
            'description' => get_the_excerpt(),
            'license'     => 'https://creativecommons.org/licenses/by/4.0/',
            'distribution' => [ [
                '@type'          => 'DataDownload',
                'encodingFormat' => 'CSV',
                'contentUrl'     => wp_get_attachment_url( 812 ),
            ] ],
        ];
    }
    return $schemas;
} );

Swap in your page slug and attachment ID. To pull a dataset out of Dataset Search later, add a robots noindex to the landing page; Google says the change can take days or weeks to show.

Warnings you can ignore, and ones you can't

Google lists two known quirks. Validators may insist that an Organization in creator needs contact details with a contactType (useful values include customer service, emergency, journalist, newsroom and public engagement), and the Rich Results Test may complain that csvw:Table is an unexpected value for mainEntity. The second one is safe to ignore. The CSVW approach for embedding tabular data is itself still in beta.

Real errors are different: a missing description, one under 50 characters, or a contentUrl that returns 404. Check syntax with the JSON-LD checker, then open every download link by hand.

Common questions

Will Dataset markup give my page a rich result in Google Search?

No. Since November 2025 Google has said it uses Dataset markup only in Dataset Search, not in regular search results.

Does a blog post with a chart count as a dataset?

Only if you publish the underlying data as something people can use, such as a downloadable file or table. A chart image alone is a weak fit.

Do I need a DOI?

No. identifier is recommended, not required. If your data has a DOI, include it because it helps Google match copies of the same dataset.

Can one page describe several datasets?

Yes, but each needs its own name and description. For a collection with parts, use hasPart on the parent and give each part the required properties.