Skip to content Skip to footer

Metadata

Metadata are structured information describing data that facilitate their discovery, understanding, utilization, and management.
They are also called "data about data" or "information about information".

Metadata help other researchers understand:

  • what a dataset contains,
  • how to use it,
  • who created it and when,
  • what access rights and licenses apply.

Thanks to metadata, datasets can be found, interpreted, and used in future research.

There are three main types of metadata:

  1. Descriptive – provide essential information for locating or identifying a dataset; they may include elements such as title, abstract, author, and keywords.  
  2. Structural – describe the relationships and dependencies between individual data sets and their elements, e.g., to facilitate navigation.  
  3. Administrative – contain information useful for managing the resource, such as creation date and method, file type, and access rights. There are several subsets of administrative data. The most commonly listed are:
    - rights management (ownership rights) and 
    - preservation (archiving and maintenance of resources). 

Metadata standard

The way metadata are structured is determined by their standard. A standard specifies both what information can be included and the requirements for its structure. Many metadata standards exist to organize how data are described. They ensure metadata have a clear structure and understandable fields – both for humans and computer programs. 

Standards can be classified as general, domain-specific, or institutional, depending on the organizations and fields that use them. 

Commonly cited general metadata standards include Dublin Core, DataCite, and the Data Documentation Initiative (DDI). They are universal across disciplines and widely used.

General information about a dataset, such as that described above, belongs to the domain of general (generic) metadata schemas. For research data, the most important schema of this kind is the DataCite standard. This schema is significant due to the popularity of the DOI identifier, which is commonly used in the research community. To obtain a DOI, data must be described according to the DataCite standard, meaning the metadata must meet certain minimum requirements.
To describe any dataset, a metadata standard must be sufficiently general, typically including only basic information (e.g., title, authors, funding). More detailed descriptions of the data or research methods are covered only broadly in these schemas. For more precise descriptions, specialized, domain-specific metadata standards are used.

The type of information that is important in describing data depends on the discipline in which the data were collected. Different fields value different data; for example, geologists require different metadata than sociologists or linguists. Therefore, selecting an appropriate domain-specific metadata standard is crucial for accurately and clearly describing data.  

For describing deposited research data, RODBUK uses the Dublin Core standard.

Where to look for an information about metadata standards used in a particular discipline?

  1. The UK’s Digital Curation Centre maintains a list of metadata standards, organized by discipline, including descriptions, extensions, tools, and examples of use.
    Recommended standards per discipline.
  2. FAIRsharing, managed by the University of Oxford, compiles information on metadata standards, repositories, databases, journal policies, and funding organizations, with a focus on FAIR principles.
  3. The Research Data Alliance website allows browsing metadata standards alphabetically and by topic, providing information on tools, examples, supporting organizations, and access via API. 

Selected metadata standards are also used by various institutions, e.g.: SDMX (ECB, Eurostat, IMF, OECD, UN), SAFE (ESA), ISO 19139 (Earth sciences), Project Open Data Metadata Schema v1.1 (U.S. federal agencies), TEI and CDWA (humanities).

Numerous tools are available for creating metadata, such as Nesstar Publisher (compliant with DDI and Dublin Core standards), and others like STATA, SPSS, eENVplus, and the Metadata Editor, designed specifically for creating metadata according to the Inspire standard.

Stopka