Data repositories

From Open Access Directory
Jump to navigation Jump to search

This list is part of the Open Access Directory.

  • This is a list of repositories and databases for open data.
  • Please annotate the entries to indicate the hosting organization, scope, licensing, and usage restrictions (if any). If a repository is open in some respects but not others, please include it with an annotation rather than exclude it.
  • If you're not sure whether a given dataset or data collection is open, post your query to Is It Open Data?
  • Related lists in OAD: Disciplinary repositories (primarily for texts, not data).
  • For news about data repositories, including some newly launched repositories not yet listed here, follow the oa.repositories.data tag of the Open Access Tracking Project.
  • See also:

Archaeology

  • Also see Social sciences.
  • Fasti Online . Subdivided in Excavation, Restauration and Survey.

Astronomy

  • Also see Physics.
  • SIMBAD Astronomical Database(perma.cc). The SIMBAD astronomical database provides basic data, cross-identifications, bibliography and measurements for astronomical objects outside the solar system.

Biology

  • Also see BCO-DMO, Marine Biology data, listed with Marine Sciences repositories.
  • Also see DataONE, Entrez databases, KNB, and PANGAEA, listed under Multidisciplinary repositories.
  • Array Express(perma.cc). Archive of Functional Genomics Data stores data from high-throughput functional genomics experiments, and provides these data for reuse to the research community.
  • BioModels(perma.cc). BioModels is a repository of mathematical models of biological and biomedical systems.
  • Cancer Imaging Archive(perma.cc). TCIA is a service which de-identifies and hosts a large archive of medical images of cancer accessible for public download.
  • The Cell: An Image Library Images of all cell types from all organisms, including intracellular structures and movies or animations demonstrating functions. This project relies upon the cell biology community to populate the library. The Cell: An Image Library™ is a freely accessible, easy-to-search, public repository of reviewed and annotated images, videos, and animations of cells from a variety of organisms, showcasing cell architecture, intracellular functionalities, and both normal and abnormal processes. The purpose of this database is to advance research, education, and training, with the ultimate goal of improving human health.
  • Database of Virulence Factors in Fungal Pathogenes (DFVF)(perma.cc). The database is expected to greatly stimulate and facilitate further studies in fungal pathogens; both experimental biologists and computational biologists can use the database and/or the predicted virulence factors to guide their search for new virulence factors and/or discovery of new pathogen-host interaction mechanisms in fungi.
  • dbGaP(perma.cc). The database of Genotypes and Phenotypes (dbGaP) was developed to archive and distribute the data and results from studies that have investigated the interaction of genotype and phenotype in Humans.
  • dpSNP(perma.cc). dbSNP contains human single nucleotide variations, microsatellites, and small-scale insertions and deletions along with publication, population frequency, molecular consequence, and genomic and RefSeq mapping information for both common variations and clinical mutations.
  • dbVar(perma.cc). dbVar is NCBI's database of human genomic Structural Variation — large variants >50 bp including insertions, deletions, duplications, inversions, mobile elements, translocations, and complex variants.
  • Dryad Dryad is an international repository of data underlying scientific and medical publications, particularly data for which no specialized repository exists. All material in Dryad is associated with a scholarly publication. Most data in the repository are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted. Dryad is a non-profit organization.
  • Electron Microscopy Data Bank (EMDB)(perma.cc). Global resource for 3-Dimensional Electron Microscopy (3DEM) structure data archiving and retrieval, news, events, software tools, data standards, validation methods, and community challenges.
  • Eukaryotic Pathogen Database Resources (EuPathDB)(perma.cc). EuPathDB (formerly ApiDB) is an integrated database covering the eukaryotic pathogens in the genera listed in the [EuPathDB (formerly ApiDB) is an integrated database covering the eukaryotic pathogens in the genera listed in our Data Summary page Data Summary] page.
  • The European Genome-phenome Archive (EGA)(perma.cc). The European Genome-phenome Archive (EGA) is a service for permanent archiving and sharing of all types of personally identifiable genetic and phenotypic data resulting from biomedical research projects.
  • Gene Expression Omnibus High-throughput functional genomic data, including all array-based applications and some high-throughput sequencing data.
  • Human Protein Atlas(perma.cc). All the data in the knowledge resource is open access to allow scientists both in academia and industry to freely access the data for exploration of the human proteome.
  • ImmPort Shared Data(perma.cc). The ImmPort project provides advanced information technology support in the archiving and exchange of scientific data for the diverse community of life science researchers supported by NIAID/DAIT and serves as a long-term, sustainable archive of research and clinical data.
  • Influenza Research Database(perma.cc). This resource contains avian and non-human mammalian influenza surveillance data, human clinical data associated with virus extracts, phenotypic characteristics of viruses isolated from extracts, and all genomic and proteomic data available in public repositories for influenza viruses.
  • KiMoSys(perma.cc). A web application for quantitative KInetic MOdels of biological SYStems.
  • MetaboLights(perma.cc). MetaboLights is a database for Metabolomics experiments and derived information.
  • MGnify(perma.cc). MGnify offers an automated pipeline for the analysis and archiving of microbiome data to help determine the taxonomic diversity and functional & metabolic potential of environmental samples.
  • Mouse Genome Informatics (MGI)(perma.cc). MGI is the international database resource for the laboratory mouse, providing integrated genetic, genomic, and biological data to facilitate the study of human health and disease.
  • National Biological Information Infrastructure A broad, collaborative program to provide increased access to data and information on the nation's biological resources. The NBII links diverse, high-quality biological databases, information products, and analytical tools maintained by NBII partners and other contributors in government agencies, academic institutions, non-government organizations, and private industry. (Note: In the President's budget for Fiscal Year 2012 the repository was terminated.)
  • NCBI Taxonomy(perma.cc). The Taxonomy Database is a curated classification and nomenclature for all of the organisms in the public sequence databases.
  • The Network Data Exchange (NDEx)(perma.cc). The NDEx Project provides an open-source framework where scientists and organizations can share, store, manipulate, and publish biological network knowledge.
  • NeuroMorpho.org(perma.cc). NeuroMorpho.Org is a centrally curated inventory of digitally reconstructed neurons associated with peer-reviewed publications.
  • OpenNEURO(perma.cc). A free and open platform for sharing MRI, MEG, EEG, iEEG, and ECoG data.
  • PeptideAtlas(perma.cc). A multi-organism, publicly accessible compendium of peptides identified in a large set of tandem mass spectrometry proteomics experiments.
  • Planet A network of European Plant Databases.
  • Protein Circular Dichroism Data Bank (PCDDB)(perma.cc). The Protein Circular Dichroism Data Bank (PCDDB) is a public repository that archives and freely distributes circular dichroism (CD) and synchrotron radiation CD (SRCD) spectral data and their associated experimental metadata.
  • ProteomeXChange(perma.cc). The ProteomeXchange Consortium was established to provide globally coordinated standard data submission and dissemination pipelines involving the main proteomics repositories, and to encourage open data policies in the field.
  • The Universal Protein Resource (UniProt) is a comprehensive resource for protein sequence and annotation data. The UniProt databases are the UniProt Knowledgebase (UniProtKB), the UniProt Reference Clusters (UniRef), and the UniProt Archive (UniParc). The UniProt Metagenomic and Environmental Sequences (UniMES) database is a repository specifically developed for metagenomic and environmental data.
  • uBio(perma.cc). uBio uses names and taxonomic intelligence to manage information about organisms.
  • VectorBase(perma.cc). A National Institute of Allergy and Infectious Diseases (NIAID) Bioinformatics Resource Center (BRC) providing genomic, phenotypic and population-centric data to the scientific community for invertebrate vectors of human pathogens.

Chemistry

  • Also see BCO-DMO, Marine Biology data, listed with Marine Sciences repositories.
  • Also see Entrez databases, listed under Multidisciplinary repositories.
  • Cambridge Structural Database The CCDC is a non-profit, charitable Institution whose objectives are the general advancement and promotion of the science of chemistry and crystallography for the public benefit.
  • ChemSpider. Hosted by the Royal Society of Chemistry.
  • ChemSynthesis. A database of chemicals and their physical properties.
  • eCrystals. From the Southampton Chemical Crystallography Group and the EPSRC UK National Crystallography Service.

Computer Science

  • CiteSeerX provides its databases of nearly 2 million documents and the associated texts and pdfs for research.
  • FreeStatistics of Irreproducible Research(perma.cc). The purpose of this project is to facilitate the creation, maintenance, and permanent storage of statistical computation objects that empower authors to publish reproducible and reusable research (in the form of a Compendium) through a series of web services.
  • GitHub keeps your public and private code available, secure, and backed up.
  • Google Code Project Hosting Project Hosting on Google Code provides a free collaborative development environment for open source projects. Each project comes with its own member controls, Subversion/Mercurial repository, issue tracker, wiki pages, and downloads section. Our project hosting service is simple, fast, reliable, and scalable, so that you can focus on your own open source development.
  • Launchpad can host your project’s source code using the Bazaar version control system. We also import over 2000 CVS, SVN, Git and Mercurial projects, so you can use Bazaar with those too.
  • ReproZip!(perma.cc). ReproZip can automatically pack your research along with all necessary data files, libraries, environment variables and options into a self-contained bundle. Then ReproZip can use that bundle to automatically set up the same original environment so anybody can reproduce the research on a different machine, without tracking down and installing the dependencies, or even having to run the same operating system.
  • SourceForge 2.7 million developers create powerful software in over 260,000 projects. Our popular directory connects more than 46 million consumers with these open source projects and serves more than 2,000,000 downloads a day. SourceForge is where open source happens.
  • SNAP Stanford Large Network Dataset Collection. The SNAP library is being actively developed since 2004 and is organically growing as a result of our research pursuits in analysis of large social and information networks. Largest network we analyzed so far using the library was the Microsoft Instant Messenger network from 2006 with 240 million nodes and 1.3 billion edges.
  • KONECT (the Koblenz Network Collection) is a project to collect large network datasets of all types in order to perform research in network science and related fields, collected by the Institute of Web Science and Technologies at the University of Koblenz–Landau.

Energy

Engineering

  • Also see Multidisciplinary repositories.
  • TRID(perma.cc). an integrated database that combines the records from TRB’s Transportation Research Information Services (TRIS) Database and the OECD’s Joint Transport Research Centre’s International Transport Research Documentation (ITRD) Database. TRID provides access to more than 1.25 million records of transportation research worldwide.

Environmental sciences

  • Also see BCO-DMO, Marine Biology data, listed with Marine Sciences repositories.
  • Also see DataONE, KNB, and PANGAEA, listed under Multidisciplinary repositories.
  • Also see Dryad, listed with Biology repositories.
  • EDI Data Portal(perma.cc). The EDI Data Portal contains environmental and ecological data packages contributed by a number of participating organizations.
  • The Marine Geoscience Data System (MGDS)(perma.cc). The Marine Geoscience Data System (MGDS) provides access to data portals for the NSF-supported Ridge 2000 and MARGINS programs, the Antarctic and Southern Ocean Data Synthesis, the Global Multi-Resolution Topography Synthesis, and Seismic Reflection Field Data Portal.
  • NERC Data Centers(perma.cc). NERC has a network of environmental data centres that provide a focal point for NERC's scientific data and information. These centres hold data from environmental scientists working in the UK and around the world.
  • Polar Data Catalogue A primarily Canadian archive of free RADARSAT imagery as well as Arctic, Antarctic, and other cryospheric data sets covering a range of disciplines, from natural sciences and policy to health and social sciences.
  • Socioeconomic Data and Applications Center (SEDAC) specializes in spatial data and services in support of human-environment research and applications, in the context of NASA’s Earth science mission and the overall U.S. Global Change Research Program.

Geology

  • Also see PANGAEA, listed under Multidisciplinary repositories.
  • IRIS (Incorporated Research Institutions for Seismology). From 100+ US universities and the National Science Foundation.

Geosciences and geospatial data

  • Also see DataONE and PANGAEA, listed under Multidisciplinary repositories.
  • EarthChem Library(perma.cc). The EarthChem Library is a data repository that archives, publishes and makes accessible data and other digital content from geoscience research (analytical data, data syntheses, models, technical reports, etc).
  • GeoNames. A database of placenames, under a CC-BY license. Founded by Marc Wick.
  • The Geosciences Network (GEON) project is a collaboration among a dozen PI institutions and a number of other partner projects, institutions, and agencies to develop cyberinfrastructure in support of an environment for integrative geoscience research. GEON is funded by the NSF Information Technology Research (ITR) program.
  • Magnetics Information Consortium (MagIC)(perma.cc). Improves research capacity in the Earth and Ocean sciences by maintaining an open community digital data archive for rock and paleomagnetic data with portals that allow users access to archive, search, visualize, download, and combine these versioned datasets.
  • National Space Science Data Center serves as the permanent archive for NASA space science mission data. "Space science" means astronomy and astrophysics, solar and space plasma physics, and planetary and lunar science. As permanent archive, NSSDC teams with NASA's discipline-specific space science "active archives" which provide access to data to researchers and, in some cases, to the general public.
  • OpenTopography(perma.cc). OpenTopography facilitates community access to high-resolution, Earth science-oriented, topography data, and related tools and resources.
  • Polar Data Catalogue A primarily Canadian archive of free RADARSAT imagery as well as Arctic, Antarctic, and other cryospheric data sets covering a range of disciplines, from natural sciences and policy to health and social sciences.
  • ShareGeo. Integrating the older GRADE (Geospatial Repository for Academic Deposit and Extraction) repository. From EDINA. (Repository discontinued.)

Linguistics

  • See the 40+ members of the Open Language Archives Community (OLAC).
  • TROLLing. Hosted by UiT. TROLLing "is designed as an archive of linguistic data and statistical code. The archive is open access, which means that all information is available to to everyone. All postings are accompanied by searchable metadata that identify the researchers, the languages and linguistic phenomena involved, the statistical methods applied, and scholarly publications based on the data (where relevant). Linguists worldwide are invited to post datasets and statistical models used in linguistic research."

Marine sciences

  • Also see DataONE and PANGAEA, listed under Multidisciplinary repositories.
  • BCO-DMO. The Biological and Chemical Oceanography Data Management Office, provides access to data sets contributed by investigators funded by the Biological and Chemical Oceanography sections of the US National Science Foundation (NSF).
  • SEAONE - Sea Open Scientific Data Publication(perma.cc). SEANOE (SEA scieNtific Open data Edition) is a publisher of scientific data in the field of marine sciences. Data published by SEANOE are available free. They can be used in accordance with the terms of the Creative Commons license selected by the author of data.

Medicine

  • Also see Entrez databases, listed under Multidisciplinary repositories.
  • caNanoLab(perma.cc) A data sharing portal designed to facilitate information sharing across the international biomedical nanotechnology research community to expedite and validate the use of nanotechnology in biomedicine.
  • [ebi.ac.uk/chembl/ CheMBL]. ChEMBL is a manually curated database of bioactive molecules with drug-like properties. It brings together chemical, bioactivity and genomic data to aid the translation of genomic information into effective new drugs.
  • Dryad Dryad is an international repository of data underlying scientific and medical publications, particularly data for which no specialized repository exists. All material in Dryad is associated with a scholarly publication. Most data in the repository are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted. Dryad is a non-profit organization.
  • FlowRepository(perma.cc). FlowRepository is a database of flow cytometry experiments where you can query and download data collected and annotated according to the MIFlowCyt standard.
  • The Health and Medical Care Archive (HMCA) is the data archive of the Robert Wood Johnson Foundation (RWJF), the largest philanthropy devoted exclusively to health and health care in the United States. Operated by the Inter-university Consortium for Political and Social Research (ICPSR) at the University of Michigan, HMCA preserves and disseminates data collected by selected research projects funded by the Foundation and facilitates secondary analyses of the data. The data collections in HMCA include surveys of health care professionals and organizations, investigations of access to medical care, surveys on substance abuse, and evaluations of innovative programs for the delivery of health care. Our goal is to increase understanding of health and health care in the United States through secondary analysis of RWJF-supported data collections.
  • MIRAGE (Middlesex medical Image Repository with a CBIR ArchivinG Environment). From JISC and Middlesex University.
  • [National Addiction & HIV Data Archive Program (NAHDAP)](perma.cc). The scope of the data housed at NAHDAP covers a wide range of legal and illicit drugs (alcohol, tobacco, marijuana, cocaine, synthetic drugs, and others) and the trajectories, patterns, and consequences of drug use as well as related predictors and outcomes.
  • Project Data Sphere, LLC, is a repository to broadly share, integrate and analyze historical, de-identified, patient-level data from academic and industry cancer Phase II-III clinical trials. Access to the Project Data Sphere platform is available to researchers affiliated with life science companies, hospitals and institutions, as well as independent researchers, at no cost and without requiring a research proposal.
  • Vivli(perma.cc). From the Center for Global Clinical Research Data. The Vivli platform includes an independent data repository, in-depth search engine and a secure research environment.

Multidisciplinary repositories

  • Also see Social Sciences.
  • Also see BCO-DMO, Marine Biology data, listed with Marine Sciences repositories.
  • DataCite (perma.cc). DataCite is a leading global non-profit organisation that provides persistent identifiers (DOIs) for research data and other research outputs.
  • Data Conservancy(perma.cc). Data Conservancy is devoted to developing institutional solutions for the challenges of data collection, preservation and re-use.
  • DataHub(perma.cc). There are thousands of datasets from financial market data and population growth to cryptocurrency prices.
  • DataONE DataONE is an international federation of data repositories containing earth observations data, including data from fields such as ecology, biology, evolution, and environmental sciences such as hydrology, oceanography, and atmospheric science. DataONE is a federation with participation from hundreds of field stations, universities, and government agencies through the DataONE Member Nodes.
  • Dryad Dryad is an international repository of data underlying scientific and medical publications, particularly data for which no specialized repository exists. All material in Dryad is associated with a scholarly publication. Most data in the repository are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted. Dryad is a non-profit organization.
  • EASY(perma.cc). EASY offers sustainable archiving of research data and access to thousands of datasets.
  • EUDAT(perma.cc). EUDAT offers heterogeneous research data management services and storage resources, supporting multiple research communities as well as individuals, through a geographically distributed, resilient network distributed across 15 European nations and data is stored alongside some of Europe’s most powerful supercomputers.
  • FigShare. Scientific publishing as it stands is an inefficient way to do science on a global scale. A lot of time and money is being wasted by groups around the world duplicating research that has already been carried out. FigShare allows you to share all of your data, negative results and unpublished figures. In doing this, other researchers will not duplicate the work, but instead may publish with your previously wasted figures, or offer collaboration opportunities and feedback on preprint figures.
  • KPBC. Regional academic repository for data in all fields. Poland
  • Microsoft Research Open Data(perma.cc). A collection of free datasets from Microsoft Research to advance state-of-the-art research in areas such as natural language processing, computer vision, and domain specific sciences. Download or copy directly to a cloud-based Data Science Virtual Machine for a seamless development experience.
  • Open Commons Consortium (OCC). The OCC is a not for profit that manages and operates cloud computing and data commons infrastructure to support scientific, medical, health care and environmental research. OCC members span the globe and include over 30 universities, companies, government agencies and national laboratories.
  • Open Science Data Cloud (OSDC). The OSDC is a data science ecosystem in which researchers can house and share their own scientific data, access complementary public datasets, build and share customized virtual machines with whatever tools necessary to analyze their data, and perform the analysis to answer their research questions. It is a one-stop shop for making scientific research faster and easier.
  • Open Science Framework (OSF) Open Science Framework serves as a scholarly commons for documentation, files, collaboration, and connecting to services for research outputs.
  • Scientific Data recommended repositories(perma.cc). Spreadsheet listing data repositories that are recommended by Scientific Data (Springer Nature) as being suitable for hosting data associated with peer-reviewed articles. Please see the repository list on Scientific Data's website for the most up to date list.
  • Tromsø Repository of Language and Linguistics (TROLLing)(perma.cc). TROLLing is designed as an archive of linguistic data and statistical code. The archive is open access, which means that all information is available to to everyone. All postings are accompanied by searchable metadata that identify the researchers, the languages and linguistic phenomena involved, the statistical methods applied, and scholarly publications based on the data (where relevant).
  • UPSpace University of Pretoria Research Repository, South Africa.
  • Webscope(perma.cc). The Yahoo Webscope Program is a reference library of interesting and scientifically useful datasets for non-commercial use by academics and other scientists. All datasets have been reviewed to conform to Yahoo's data protection standards, including strict controls on privacy.
  • Zenodo(perma.cc). All research outputs from across all fields of research.

Physics

  • Also see Astronomy.
  • HEP Data The data comprise total and differential cross sections, structure functions, fragmentation functions, distribuitions of jet measures, polarisations, etc... from a wide range of interactions.
  • Nist Atomic Spectra Database The Atomic Spectra Database (ASD) contains data for radiative transitions and energy levels in atoms and atomic ions. Data are included for observed transitions of 99 elements and energy levels of 56 elements.

Social sciences

  • Also see Multidisciplinary repositories.
  • Archeology Data Service(perma.cc). Heritage data, with over 20 years of experience supporting research, learning and teaching with free, high quality and dependable digital resources.
  • Databrary A repository for sharing and reusing research video data and related metadata in the developmental and learning sciences. Hosted at New York University with support from The Pennsylvania State University.
  • European Nucleotide Archive(perma.cc). The European Nucleotide Archive (ENA) provides a comprehensive record of the world's nucleotide sequencing information, covering raw sequencing data, sequence assembly information and functional annotation.
  • ICPSR (Inter-University Consortium for Political and Social Research). At the University of Michigan.
  • openICPRS(perma.cc). openICPSR is a great place to share and store your social and behavioral science research data. Your data will be preserved as-is and be available to data users at no cost.
  • Qualitative Data Repository(perma.cc). QDR curates, stores, preserves, publishes, and enables the download of digital data generated through qualitative and multi-method research in the social sciences.