BioSamples

Overview

BioSamples is a database, run by EMBL-EBI, that stores information about the biological samples used in research: the tissue, cell line, or participant that a dataset was generated from, with around 60 million samples recorded as of 2025. Its purpose in neuroscience genomics is to keep that sample information in one place, because a single study’s data is usually split across several archives, with the raw sequence held in one and the expression or variant data in another. BioSamples gives each sample one record with a permanent identifier, called an accession, that every one of those archives can point back to. An accession issued by BioSamples begins with the letters SAMEA followed by a number. Each record describes the sample through structured attributes such as species, tissue, and age, written with shared vocabularies so that the same term means the same thing across studies. Every sample submitted to ENA is automatically given a BioSamples record, so the genomic, transcriptomic, and phenotypic data taken from one brain tissue sample or cohort participant all trace back to the same entry.

Connections

  • operatedBy: EMBL
  • isPartOf: ELIXIR
  • relatedTo: ENA (samples registered in ENA are mirrored in BioSamples with a SAMEA accession)

Resources