Sharing your data

Sharing a neuroscience dataset openly requires decisions at each stage of the research cycle: what consent to obtain, what format to use, where to deposit, and how to register the result so others can find and cite it. This perspective walks through those stages in order. For modality-specific standards and repositories, see the relevant domain perspective: Neuroimaging, Electrophysiology, Genomics, Bioimaging, Cognition, or Computational. For funder mandates and national platforms, see the relevant geographic perspective.

Before you collect

Open Brain Consent provides GDPR-compatible model informed consent forms designed specifically for neuroimaging and electrophysiology data from human participants. The forms explicitly allow open data sharing and have been reviewed against European data protection requirements. Consent obtained without an open-sharing clause cannot be retroactively revised, which makes this the step with the highest downstream cost if skipped.

What counts as sensitive data

Whether a dataset can be shared openly or must go to controlled access depends on whether the applicable law treats its contents as sensitive. That threshold differs by jurisdiction. Establish which regime applies, usually that of the institution holding the data and of the participants, before promising open access in a consent form or data management plan. The same dataset can be openly shareable under one regime and restricted under another.

The three regimes a reader is most likely to meet define the protected category differently:

  • European Union (GDPR). GDPR sets out a closed list of “special categories” in Article 9. It covers health and genetic data, biometric data used for unique identification, racial or ethnic origin, political opinions, religious or philosophical beliefs, trade union membership, and sex life or sexual orientation. Most identifiable neuroscience data falls inside it: clinical records and genomic data are special-category by definition, and identifiable neuroimaging can be too. Such data may be processed only under a derogation. For research, the route is the scientific research provision (Article 9(2)(j)) with the safeguards of Article 89, such as pseudonymisation. The list is fixed, so membership is settled by category rather than case by case.
  • United States. There is no single federal law equivalent to GDPR. Protection is sectoral. HIPAA covers health information held by defined covered entities (providers, plans, clearinghouses) and their business associates. Other federal statutes cover finance or education, and a growing set of state laws (such as California’s) add their own definitions, several treating biometric or health data as sensitive. For research the operative question is rarely the category. It is whether the data has been de-identified to the HIPAA standard, because US funder policy, notably the NIH Data Management and Sharing Policy, requires this before deposit. HIPAA does not regulate research archives directly.
  • Canada (PIPEDA). Federal law gives no statutory list. Sensitivity is judged case by case, with safeguards proportionate to it. In practice the regulator and the courts treat health, genetic, and biometric data as generally sensitive, and the privacy commissioner’s guidance mirrors the GDPR categories. Jurisdiction is split between federal and provincial law, so provincial health-privacy statutes (such as Ontario’s PHIPA) and research-ethics rules (TCPS2) apply alongside the federal framework.

De-identification is not anonymisation

Clearing a de-identification standard and being free to share openly are not the same thing. The two main regimes set the bar in different places, so a dataset can clear one and not the other.

Under HIPAA, data is de-identified once the Safe Harbor identifiers are removed, or an expert certifies the re-identification risk is very small. It is then no longer protected health information, and US rules permit sharing. Safe Harbor is a checklist: remove the listed fields and the test is met. It does not weigh indirect identifiers, so a rare diagnosis plus a few retained fields can still single out a person. HIPAA does not use the word anonymisation.

Under GDPR the test is an outcome, not a checklist. Data is anonymised only when re-identification is reasonably impossible by any available means, including linkage to outside datasets. Only then does it leave GDPR and become freely shareable. Data stripped of direct identifiers but still carrying re-identification risk is pseudonymised, which remains personal data and stays regulated. HIPAA de-identification generally maps to GDPR pseudonymisation, not anonymisation, because the checklist leaves linkage risk that GDPR’s outcome test counts. The reasoned position is to treat HIPAA-de-identified data as pseudonymised, and still regulated, under GDPR.

The effect on a sharing decision is direct. The same dataset can be openly releasable under US rules but, under GDPR, releasable only under controlled access, because it remains personal data. Anonymisation strong enough to pass the all-available-means test usually removes enough detail to undercut the research use, so the realistic route for sensitive neuroscience data is not open release but controlled access under a data use agreement. This is sharpest for inherently identifying data such as whole-genome sequences and full-face structural scans, where no field removal makes the data non-identifying. For these, controlled access is the norm in every jurisdiction. The Health and Genomics perspectives list the controlled-access repositories, and HIPAA and GDPR carry the underlying definitions.

Study registration

Pre-registering a study design before data collection reduces publication bias and is increasingly recognised by journals and funders. OSF is the standard platform for pre-registration in non-clinical neuroscience research. For clinical trials, registration before first patient enrolment is a legal requirement: ClinicalTrials.gov is the primary global registry, and CTIS is mandatory for all trials conducted under EU clinical trials legislation.

Data management plan

Most research funders require a data management plan specifying where data will be stored, how it will be described, and under what conditions it will be shared. Making format and repository choices before writing the plan makes it straightforward to complete. Common DMP tools include OPIDoR (France, ANR-aligned), DMPTool (USA, NIH-aligned), DMP Online (UK, UKRI/Wellcome-aligned), and the Data Stewardship Wizard (EU/ELIXIR, with maDMP export). RDMkit, maintained by ELIXIR, provides comprehensive RDM guidance covering the full data lifecycle, navigable by life science domain and by country. For funder-specific requirements and pre-registration as standalone topics, see Study Registration and Data Management.

Formatting your data

Every research modality has a community-adopted format standard. Using it is the single most important step for repository acceptance and interoperability downstream. Validate against the format standard before deposit: most domain repositories run validation on upload and reject non-compliant datasets. The domain perspectives describe each standard in depth. The reference below covers the working choices.

ModalityStandard formatSee
MRI, fMRI, PETBIDS over NIfTINeuroimaging
EEG, MEG, iEEGBIDS with EDF or BrainVision, plus HED for eventsElectrophysiology
Invasive neurophysiologyNWB, plus HED for eventsElectrophysiology
Genomics (raw to processed)FASTQSAM-BAM-CRAMVCFGenomics
Single-cell genomicsh5ad (AnnData / Seurat)Genomics
BioimagingOME File Formats (OME-TIFF or OME-Zarr)Bioimaging
Behavioural and cognitiveBIDS with HED event sidecar filesCognition
Computational modelsNeuroML (models), SWC (morphologies)Computational
Clinical trial dataCDISC (SDTM, ADaM)Clinical Trials
Health dataOMOP CDM harmonisation applied at source by the institution holding the dataHealth

Choosing a repository

The right repository depends on data modality, access policy, and any funder or institutional requirements. Domain perspectives list the full repository landscape for each modality. The table below gives a broad overview with examples. For country-specific deposit requirements and recommended national platforms, see the geographic perspective for your region.

ModalityOpen accessControlled accessSee
NeuroimagingOpenNeuroPublic-nEUro (EU), NIMH Data Archive (US)Neuroimaging
ElectrophysiologyDANDI ArchiveG-Node (EU)Electrophysiology
GenomicsENA / SRA / DDBJEGA (EU), dbGaP (US)Genomics
Single-cell genomicsCELLxGENE, NeMO ArchiveEGA (EU), dbGaP (US)Genomics
BioimagingBioImage ArchiveBioimaging
Behavioural / cognitiveOpenNeuro (with recordings), OSF or Zenodo (standalone)Cognition
Computational modelsModelDB, OpenSourceBrainComputational
Clinical trial dataResults via ClinicalTrials.govSee Clinical TrialsClinical Trials
Health dataSee HealthEGA (EU), dbGaP (US), and national platforms (see geographic perspective)Health
Any modalityZenodo
Discover domain repositoriesFAIRsharing, re3data

When no domain-specific or institutional repository fits, Zenodo and Figshare are general-purpose options that accept any file type and assign a DataCite DOI on deposit. Zenodo, operated by OpenAIRE and CERN, is the recommended destination for EU-funded projects without a domain repository.

Depositing and registering

Deposit the dataset with complete metadata, including ORCID identifiers for all contributors. Most repositories assign a DOI via DataCite automatically on submission, making the dataset citable from the moment of deposit.

After depositing, register the dataset in FAIRsharing to link it to the standards it uses and make it findable alongside related repositories and policies. FAIR-Checker can then assess the FAIRness of the deposited resource by its URL or DOI, identifying specific gaps in metadata completeness, structure, and interoperability against recognised standards. If the study was pre-registered, link the deposit record to the pre-registration entry in OSF or the relevant clinical trials registry. Include the repository DOI in a data availability statement in your paper. A statement saying only “data available on request” requires no actual deposit and routinely results in low retrieval rates. Publishers aligned with cOAlition S increasingly require a verified deposit DOI rather than a request-based statement.

For the provenance and reproducibility layer (open code, workflow tracking, and re-executable pipelines alongside your data), see Reproducibility.