Shared Datasets
Shared datasets are datasets that are stored centrally and made available to multiple ML Cloud users or research groups.
Using a shared copy is an important part of responsible research data management. It reduces unnecessary duplication, helps users work with consistent dataset versions, and makes commonly used data easier to document and maintain.
Available Shared Datasets
Shared datasets are currently provided through the high-performance filesystems of the two ML Cloud clusters. Shared datasets are currently provided through the high-performance filesystems of the two ML Cloud clusters:
Check these lists before downloading a large dataset. The dataset you need may already be available.
Do You Want A New Dataset?
If you would like an additional dataset installed for general use, please contact us through the ticketing system with all relevant information(name, description, restrictions, etc.)
Why use shared datasets?
Large datasets are often downloaded or copied independently by different users. This can consume substantial storage space and may result in several incomplete, outdated, or differently processed versions of the same dataset.
Using a centrally managed copy helps to:
- reduce unnecessary data duplication
- improve storage efficiency
- make datasets easier to find and reuse
- support consistent dataset versions
- preserve information about the dataset source and licence
- improve the reproducibility of research workflows
Working with shared datasets
Shared datasets should normally be treated as read-only source data.
Where possible, use the centrally available dataset directly rather than creating another complete copy. Store any project-specific files separately, including:
- processed or converted data
- extracted subsets
- annotations
- train, validation, or test splits
- temporary files
- model outputs and checkpoints
- analysis results
This separation helps protect the original dataset and makes it clearer which files were obtained from the source and which were created during your research.
When is a copy necessary?
A separate copy may be appropriate when:
- the original files must be modified
- only a project-specific subset is needed
- a fixed version must be retained for reproducibility
- the workflow cannot use the shared location directly
When creating a copy or derived dataset, document its relationship to the shared source and the changes that were made.
Avoid creating complete copies only for convenience, particularly for large datasets.
Licensing and access conditions
Access to a dataset does not automatically mean that it can be used or redistributed without restrictions.
Before using a shared dataset, check:
- its licence and terms of use
- whether citation or acknowledgement is required
- whether access is restricted
- whether derived data may be shared
- whether registration with the original provider is required
Users remain responsible for complying with the conditions of the original dataset provider.
Best practices
When using shared datasets:
- check whether the dataset is already available
- use the shared copy directly where possible
- do not modify the original shared files
- store derived data and outputs separately
- document the dataset version and source
- follow licence and access restrictions
- report missing, outdated, or damaged files through a support ticket
Created: June 21, 2024