Skip to content

Step 1: Prepare your data

Before depositing your data in DataverseNO, spend some time organizing and documenting your dataset. Doing so will make it easier for others to understand, reuse, and cite your data, and will help ensure a smooth publication process.

Most datasets can be prepared by following a few simple steps:

You do not need to get everything perfect. The purpose of these guidelines is to help you prepare data that can be understood and reused by others. If needed, support staff can provide additional guidance during the curation process before publication.

If you are unsure what applies to your dataset, please contact your local user support. We are happy to help.

Organize your files

Clear file organization makes it easier for collaborators, curators, and future users to navigate your dataset.

Good practice for file names

  • Avoid spaces and special characters (like \ / ? : * ” > < | : # % ” { } | ^ [ ] ` ~ æ ø å ä ö).

  • Use the date format YYYY-MM-DD.

  • Keep file names reasonably short (<25 characters).

  • Use descriptive file names.

  • Use consistent naming conventions.

Example: Good file names

  • 00_README.txt
  • survey_data_2025-08.csv
  • species_observations_2024-06-27.tsv
  • interview_metadata.xlsx

Example: Bad file names

  • data 1.xlsx
  • group Ø-Å final final NEW.xlsx
  • test.docx
  • untitled.csv

Spreadsheets and tabular data

For spreadsheets and tabular files, we recommend:

  • One table per file.

  • One row per observation.

  • One column per variable.

  • One value per cell.

  • Variable names without spaces and special characters (like \ / ? : * ” > < | : # % ” { } | ^ [ ] ` ~ æ ø å ä ö).

  • Useing the date format YYYY-MM-DD.

For more detailed guidance, see chapter Data Organisation in Spreadsheets in The Turing Way handbook to reproducible, ethical and collaborative data science.

Choose suitable file formats

Why do file formats matter?

Some file formats are easier to preserve and reuse than others. DataverseNO therefore recommends a number of preferred file formats for long-term access and reuse.

However, many datasets can still be published in their original formats. The use of a non-preferred format does not automatically prevent publication.

If your data were originally created in a non-preferred file format, we often recommend uploading both the preferred file format and the original file format. The preferred file format supports long-term preservation and reuse, while the original file format may be easier for some users to inspect or work with in the short term.

Preferred file formats

Examples include:

Data type Preferred formats
Text TXT, PDF/A
Tabular data TSV, CSV
Images TIFF, PNG, JPEG
Audio WAV, AIFF, FLAC
Video MP4
Markup XML, HTML
Statistical data R, SPSS syntax, STATA syntax
Software code Python, MATLAB, plain-text source code

For the complete list including guidance, see the DataverseNO file formats overview.

Need help converting files?

Guidance on how to convert documents, spreadsheets, images, audio files, video files, and other data types into preferred file formats is available in the DataverseNO file formats overview.

Original and converted file formats

Often, data may be provided both in a preferred file format and in the original file format from which it was derived.

Example:

experiment_01.csv
experiment_01.xlsx

In this example:

  • experiment_01.csv is the preferred preservation format.

  • experiment_01.xlsx is the original working format.

Providing both versions can help support both long-term preservation and immediate reuse.

Keep in mind

If you upload both a preferred file format and an original file format, the file names should be identical except for the file extension.

Describe your data

Good documentation makes data easier to find, understand, and reuse.

In DataverseNO, datasets are documented in two complementary ways:

  • A README file that you prepare before depositing your dataset.

  • Metadata that you enter when creating the dataset in DataverseNO.

This section focuses on preparing a README file and other documentation before deposit. Information about metadata is provided in the second step of the deposit workflow: Step 2: Deposit your data.

The most important step: Create a README file

A README file is a guide to your dataset. It explains what the data contain, how they were created, how the files are organized, and what someone needs to know to understand and reuse them.

Providing a README file is required before a dataset can be published in DataverseNO.

The file should, at a minimum, contain:

  • Dataset title and contact information.

  • Methodological information.

  • Overview of files and folders.

  • Explanations of variables, abbreviations, codes, or terminology.

  • Information about terms of reuse and licensing.

We recommend using the DataverseNO template:

DataverseNO README File Template Example 1 (Life Sciences) Example 2 (Social Sciences)

Keep in mind

A well-written README file is often the single most important factor enabling others to understand and reuse your data.

Additional documentation

Depending on the nature of your dataset, it may be helpful to include additional documentation alongside the README file and refer to it in the README file where relevant.

Examples include:

  • Data collection protocols.

  • Analysis scripts.

  • Codebooks.

  • Survey instruments or interview guides.

  • Processing workflows.

  • Laboratory procedures.

  • Documentation of rights and permissions.

The more specialized your dataset is, the more important such documentation often becomes.

Check file and dataset size

To ensure smooth uploading, curation, and reuse, please note the following recommendations:

  • Individual files should preferably not exceed 100 GB.

  • A single upload should preferably not exceed 200 GB.

  • A dataset should preferably not exceed 5 TB.

  • A dataset should preferably not contain more than 300 files.

Larger datasets

If your files or dataset exceed these recommendations, contact your local user support before depositing your data. Large datasets can often be accommodated.

Need help?

If you are unsure how to prepare your dataset, contact your local user support. We are happy to help.

Ready to continue?

If your files are organized, documented, and ready to share, you are ready for the next step.

Step 2: Deposit your data