17 Data storage
17.1 Learning objective
At the end of the unit, you will be familiar with some key concepts of research data management and know the key aspects for preparing your data for long-term storage and reuse.
17.2 Prior knowledge
No prior knowledge is expected for this unit.
17.3 Material
This unit combines text with several videos and a matching game.
17.4 Learning content
Proper storage and management of (research) data is as important as the data itself and a fundamental aspect of good research practice. Otherwise, the data might not be comprehensible and reusable by other researchers (including yourself several years later), as illustrated in this video:
The video also illustrates that this includes the metadata, i.e. data describing the data. In most of the cases, it is only with the metadata that the data become comprehensible. Examples for typical metadata for Pb isotope data are the analytical protocol used to prepare the sample for analysis, location, and information about the geological or archaeological context such as the ore or material type. Ideally, strategies and workflows for proper data management are planned before the start of a project. This is often done by writing a data management plan which includes, among others, details about which data is collected, how they are handled, and how they can be accessed by whom during the project and after. This information is also requested by an increasing number of funding organisations either as a separate data management plan or comparable chapters in the proposal.
17.4.1 Data life cycle
In general, the major steps in handling data can be modelled along the so-called data life cycle. The data life cycle describes the different stages from the creation to the publication of data and their reuse in subsequent projects. While versions with different levels of detail exist (Ball 2012), the following video explains the basic stages along a typical use case for Pb isotope data:
17.4.2 FAIR data principles
If you delve into research data management, you will sooner than later come across the FAIR data principles. FAIR is an acronym for Findable, Accessible, Interoperable, and Reusable. Defined by (Wilkinson et al. 2016), these principles describe what it needs to optimise data and its metadata to make them available for future research in the best way possible. The emphasis of the FAIR principles is on machine-actionability, meaning that the data can be processed by algorithms to e.g., download and combine datasets from different sources, but following these principles will also optimise your data for human readers. An increasing number of funders such as the European Commission require that data created in projects funded by them adhere to the FAIR data principles.
FAIR data does not mean open data but follows the principle of “as open as possible, as closed as necessary”. Of course, this varies depending on the type of data and for example, health-related data on the level of an individual will not be disclosed. However, it is still important to know that such data exists and what it covers. And that’s what the FAIR data principles are about: making sure that there is sufficient information about which data exists and how the data is organised (Interoperability), where to find it (Findability), and under which conditions it can be accessed (Accessibility) and reused (Reusability).
A very concise overview about what exactly the FAIR principles include is given in this short video from the Sustainable Digital Scholarship :
In the following, the FAIR principles will be briefly summarised. Many webpages and videos provide more in-depth information about them and how to implement them, such as How to FAIR (Deutz et al. 2020), go-fair.org or the video The FAIR principles explained by the YouTube channel FAIR Enough.
Findable means that data can be easily found. Most importantly, this includes that the data and its metadata are assigned a persistent and globally unique identifier. Usually, this is a link which is guaranteed to always resolve a specific dataset, even if the actual URL changes and only to this dataset and nothing else. Examples for such special identifiers are the Digital Object Identifier (DOI) and ORCID iD. In addition, the metadata are searchable, i.e. they can be accessed by search engines.
Accessible means that the data can be accessed through standardised communication protocols, such as application programming interfaces (API) and these protocols are open, free to use and, if needed, allows for authentication. It further means that metadata are accessible even if the data itself are not or not anymore – as mentioned before FAIR data are not necessarily open data.
Interoperable means in fact that the data are understandable for both humans and machines. This requires that metadata are described with vocabularies which itself are following the FAIR principles, that the metadata include references to other (meta)data and that they are structured in a way that can be easily accessed by machines. This is required to allow machines making inferences, such as knowing that you are looking for data about “bronze” as metal and not the colour or that when searching for weights in gram data reported in milligram or kilogram should be included as well.
Reusable means that clear information is provided about what you are allowed to do with the data, i.e. a usage licence, and where the data are coming from, i.e. detailed information about their provenance. In addition, this principle also includes that the data and metadata are following standards relevant for their respective domain to facilitate their combination with other datasets. This does not only refer to e.g. the vocabularies used to ensure interoperability but also to the file types the data is stored as and using common templates to describe and document your data.
17.4.3 Ethical handling of data
Ethical handling of archaeological remains such as human remains or items related to descendant communities is a very important aspect of good archaeological practice. This likewise applies to data obtained from such materials (Gupta et al. 2023). While lead isotope data were already obtained from human remains Erel et al. (2021), the predominant aspect is probably data obtained from objects (archaeological and geological) found on the lands of or (potentially) owned by descendant communities.
Descendant communities have good reasons for disclosing such data. For example, publishing data from ores might raise unwanted interest by mining companies even if data were not obtained in such a context. Likewise, publishing about items might attract looters. It is also important to keep in mind that (invasive) sampling of metal items might go against their belief system for various reasons or is not wanted to keep the objects intact.
However, many of them are open for collaboration on eye level. It is therefore important to investigate beforehand if the items you want to analyse could be related to a descendant community and to get in touch with them. To raise awareness of potential conflicts between the interest of descendant communities and researchers in the way indigenous data are collected, used, published, and reused, the CARE principles for Indigenous Data Governance (Carroll et al. 2020) were drafted. CARE stands for Collective Benefit, Authority to Control, Responsibility, Ethics. The FAIR and CARE principles are complementary but not fully compatible. This video from the TETRARCH project on Vimeo explains in more detail how both sets of principles can be implemented in archaeological research:
In addition, the Traditional Knowledge (TK) Labels are an initiative for descendant communities and other local organisations to facilitate understanding of the importance of the data for users outside of the respective communities. They provide a clear indication who owns the data and who should be contacted for reuse, traditional protocols associated with the material, and under which conditions it can be used or reused.
17.4.4 Making data fit for publication and archiving
In the context of digital data, long-term preservation refers to a time-period where the data are likely to be affected by technological change, usually assumed to be around 10 years, and the durability of the hardware the data is stored on (Lunt et al. 2024; TexasDigitalArchive-n-d?). Keeping data available for longer than 10 years is usually referred to as long-term archiving. It is safe to assume that lead isotope data must be prepared for long-term archiving. For example, Pb isotope data measured in the 1980ies (Stos-Gale and Gale 2009) are still regularly used for raw material provenancing and they will be for the foreseeable future (e.g. Tomczyk 2022).
Long-term archiving of digital data comes with many challenges which were identified for example in the chapter Digital preservation briefing of Digital Preservation Coalition (2015) and by The Consultative Committee for Space Data Systems (2012). Among them, three aspects are of particular concern for researchers and should be addressed here:
- Where to store and archive the data?
- In which file format should the data be used?
- How to make the data findable and reusable?
17.4.4.1 Where to store the data
The only suitable infrastructure for long-term preservation and archiving of data are data repositories and archives. Data repositories are infrastructures whose operators commit themselves to keep the data available in the long-term, ideally including a strategy for long-term archiving. Their upkeep is financially backed up by the institutions running them. Repositories differ in the topics or institutions they cover, the services they offer and their operational model. An example for a very generic, free-to-use repository is Zenodo. Everyone with a suitable user account can deposit their data in Zenodo, however no data curation services are provided. Consequently, it is up to you, the data provider, to make sure that data is properly described and optimised for reuse. In contrast, domain-specific repositories accept only data related to a specific domain but usually provide support in the organisation and description of the data. Here, your data is usually curated by domain experts, which ensures that they are described and structured according to the domain’s standards. While some of them do not charge fees (e.g., GFZ Data Services), others do (e.g. OpenContext, Archaeological Data Service). In addition, many research institutions and universities nowadays have their own repositories, available to the respective staff (i.e., institutional repositories).
Some repositories are certified, with the CoreTrustSeal being currently the most widespread international certificate for data repositories. One evaluation criterion of the CoreTrustSeal is the existence of strategies for the long-term preservation and archiving of data the repository holds. However, it must be kept in mind that these certifications require substantial effort from the repositories and it may be that repositories conform already to these criteria but cannot spare the resources for the certification process.
Domain-specific repositories are recommended for long-term preservation and archiving because they are the experts with regards to which metadata standards are established in the community and how data should be described and optimised for reuse by your peers. Most of the generic and domain-specific repositories and key information about their scope and services can be found in re3data. Especially domain-specific repositories should ideally be already contacted in the design stage of a project to discuss requirements for the data from the different sides (esp. data provider, funder, repository) and to learn about fees that may apply.
17.4.4.2 File formats
Choosing the right file format is decisive in keeping the data accessible. Many software programmes, even open-source ones, store data in a proprietary file format by default. Reading them requires access to the respective software or software that can convert it into another file format. Software and file formats get out of use quickly, especially if not maintained anymore and/or restricted to specific (versions of) operating systems. Therefore, file formats suitable for long-term preservation
- Are commonly used
- Can be read by a wide range of programmes
- Have well-documented technical specifications
- Are open
- Are non-proprietary
Many lists with suitable formats for different media types exist (e..g, from the Library of Congress, Smithsonian Institution, and the Swedish National Data Service). Preferences may vary among the repositories and checking them already in the design stage of a project helps to optimise data handling workflows from the onset. The quiz below allows you to learn about typical file formats for different file types.
17.4.4.3 Making data findable and reusable
You can publish to the highest possible standards but if no one knows your data exists or, if they know it exists, your data are not licensed for being reused. Therefore, making your data findable and reusable is the third important aspect. Findability and Reusability are two of the four major aspects in the FAIR data principles. Most notable in the context of long-term preservation is that the data have a persistent identifier such as a DOI. Such an identifier is usually assigned by the repository when it makes the data available. With the help of this identifier, and the metadata stored with it, the data can be easily found on the internet through e.g., search engines. In addition, it is only with clear (and ideally machine-readable) information about the usage rights that data can be safely reused by others. Usually, this is done by licensing the data publication under a Creative Commons licence. Creative Commons licences aim to bridge the different national copyright laws by providing a modular system of restrictions on how the data can be reused. The most common licence for data publications and open access articles is probably Creative Commons Attribution, which is roughly comparable to the obligation for including a citation.
17.5 Self check
Now you can provide answers to the following questions:
- Why is it important to properly organise and describe data?
- What are metadata and why are they so important?
- To which steps of the data life cycle are the FAIR data principles contributing the most?
- Which file format should you use to preserve a spreadsheet for long-term storage?
- What are data repositories and where can you find the ones suiting your data best?
- Why are globally unique and persistent identifiers so important for the data ecosystem?
17.6 Further reading
- Brinken H, Hauss J, Rücknagel J (2021) The 101 of Creative Commons licenses. Video series (5 episodes). https://av.tib.eu/series/1786.
- O’Brien M, Duerr R, Taitingfong R, Martinez A, Vera L, Jennings LL, Downs RR, Antognoli E, Brink TT, Halmai NB, David-Chavez D, Carroll SR, Hudson M, Buttigieg PL (2024) Earth Science Data Repositories: Implementing the CARE Principles. Data Science Journal 23:95. https://doi.org/10.5334/dsj-2024-037
- Strasser C, Cook R, Michener W, Budden A (2012) Primer on Data Management: What you always wanted to know. https://doi.org/10.5060/D2251G48