Step 1: Prepare your data¶
Before depositing your data in DataverseNO, spend some time organizing and documenting your dataset. Doing so will make it easier for others to understand, reuse, and cite your data, and will help ensure a smooth publication process.
Most datasets can be prepared by following a few simple steps:
-
Organize your files so that they are easy to navigate and understand.
-
Choose suitable file formats that support long-term access and reuse.
-
Describe your data using a README file and other relevant documentation.
-
Check file and dataset size before uploading your files.
You do not need to get everything perfect. The purpose of these guidelines is to help you prepare data that can be understood and reused by others. If needed, support staff can provide additional guidance during the curation process before publication.
If you are unsure what applies to your dataset, please contact your local user support. We are happy to help.
Organize your files¶
Clear file organization makes it easier for collaborators, curators, and future users to navigate your dataset.
Good practice for file names¶
-
Avoid spaces and special characters (like \ / ? : * ” > < | : # % ” { } | ^ [ ] ` ~ æ ø å ä ö).
-
Use the date format YYYY-MM-DD.
-
Keep file names reasonably short (<25 characters).
-
Use descriptive file names.
-
Use consistent naming conventions.
Example: Good file names
- 00_README.txt
- survey_data_2025-08.csv
- species_observations_2024-06-27.tsv
- interview_metadata.xlsx
Example: Bad file names
- data 1.xlsx
- group Ø-Å final final NEW.xlsx
- test.docx
- untitled.csv
Spreadsheets and tabular data¶
For spreadsheets and tabular files, we recommend:
-
One table per file.
-
One row per observation.
-
One column per variable.
-
One value per cell.
-
Variable names without spaces and special characters (like \ / ? : * ” > < | : # % ” { } | ^ [ ] ` ~ æ ø å ä ö).
-
Useing the date format YYYY-MM-DD.
For more detailed guidance, see chapter Data Organisation in Spreadsheets in The Turing Way handbook to reproducible, ethical and collaborative data science.
Choose suitable file formats¶
Why do file formats matter?¶
Some file formats are easier to preserve and reuse than others. DataverseNO therefore recommends a number of preferred file formats for long-term access and reuse.
However, many datasets can still be published in their original formats. The use of a non-preferred format does not automatically prevent publication.
If your data were originally created in a non-preferred file format, we often recommend uploading both the preferred file format and the original file format. The preferred file format supports long-term preservation and reuse, while the original file format may be easier for some users to inspect or work with in the short term.
Preferred file formats¶
Examples include:
| Data type | Preferred formats |
|---|---|
| Text | TXT, PDF/A |
| Tabular data | TSV, CSV |
| Images | TIFF, PNG, JPEG |
| Audio | WAV, AIFF, FLAC |
| Video | MP4 |
| Markup | XML, HTML |
| Statistical data | R, SPSS syntax, STATA syntax |
| Software code | Python, MATLAB, plain-text source code |
For the complete list including guidance, see the DataverseNO file formats overview.
Need help converting files?¶
Guidance on how to convert documents, spreadsheets, images, audio files, video files, and other data types into preferred file formats is available in the DataverseNO file formats overview.
Original and converted file formats¶
Often, data may be provided both in a preferred file format and in the original file format from which it was derived.
Example:
experiment_01.csv
experiment_01.xlsx
In this example:
-
experiment_01.csv is the preferred preservation format.
-
experiment_01.xlsx is the original working format.
Providing both versions can help support both long-term preservation and immediate reuse.
Keep in mind
If you upload both a preferred file format and an original file format, the file names should be identical except for the file extension.
Describe your data¶
Good documentation makes data easier to find, understand, and reuse.
In DataverseNO, datasets are documented in two complementary ways:
-
A README file that you prepare before depositing your dataset.
-
Metadata that you enter when creating the dataset in DataverseNO.
This section focuses on preparing a README file and other documentation before deposit. Information about metadata is provided in the second step of the deposit workflow: Step 2: Deposit your data.
The most important step: Create a README file¶
A README file is a guide to your dataset. It explains what the data contain, how they were created, how the files are organized, and what someone needs to know to understand and reuse them.
Providing a README file is required before a dataset can be published in DataverseNO.
The file should, at a minimum, contain:
-
Dataset title and contact information.
-
Methodological information.
-
Overview of files and folders.
-
Explanations of variables, abbreviations, codes, or terminology.
-
Information about terms of reuse and licensing.
We recommend using the DataverseNO template:
DataverseNO README File Template Example 1 (Life Sciences) Example 2 (Social Sciences)
Keep in mind
A well-written README file is often the single most important factor enabling others to understand and reuse your data.
Additional documentation¶
Depending on the nature of your dataset, it may be helpful to include additional documentation alongside the README file and refer to it in the README file where relevant.
Examples include:
-
Data collection protocols.
-
Analysis scripts.
-
Codebooks.
-
Survey instruments or interview guides.
-
Processing workflows.
-
Laboratory procedures.
-
Documentation of rights and permissions.
The more specialized your dataset is, the more important such documentation often becomes.
Check file and dataset size¶
To ensure smooth uploading, curation, and reuse, please note the following recommendations:
-
Individual files should preferably not exceed 100 GB.
-
A single upload should preferably not exceed 200 GB.
-
A dataset should preferably not exceed 5 TB.
-
A dataset should preferably not contain more than 300 files.
Larger datasets¶
If your files or dataset exceed these recommendations, contact your local user support before depositing your data. Large datasets can often be accommodated.
Need help?¶
If you are unsure how to prepare your dataset, contact your local user support. We are happy to help.
Ready to continue?¶
If your files are organized, documented, and ready to share, you are ready for the next step.