Data repositories
Jump to navigation
Jump to search
This list is part of the Open Access Directory.
- This is a list of repositories and databases for open data.
- Please annotate the entries to indicate the hosting organization, scope, licensing, and usage restrictions (if any). If a repository is open in some respects but not others, please include it with an annotation rather than exclude it.
- If you're not sure whether a given dataset or data collection is open, post your query to Is It Open Data?
- Related lists in OAD: Disciplinary repositories (primarily for texts, not data).
- For news about data repositories, including some newly launched repositories not yet listed here, follow the oa.repositories.data tag of the Open Access Tracking Project.
- See also:
- Re3data.org. The re3data.org project intends to create a global registry of research data repositories.
- Recommended Data Repositories from Nature.
- FAIRsharing. The FAIRsharing project compares data repositories for their compliance with the FAIR principles and journal data-sharing policies.
Archaeology
- Also see Social sciences.
- Fasti Online . Subdivided in Excavation, Restauration and Survey.
- Open Context. From the Alexandria Archive Institute.
Astronomy
- Also see Physics.
- Astrophysics Data System. From the Smithsonian Astrophysical Observatory (SAO) and National Aeronautics and Space Administration (NASA).
- National Space Science Data Center. From the US National Aeronautics and Space Administration (NASA).
- SIMBAD Astronomical Database(perma.cc). The SIMBAD astronomical database provides basic data, cross-identifications, bibliography and measurements for astronomical objects outside the solar system.
Biology
- Also see BCO-DMO, Marine Biology data, listed with Marine Sciences repositories.
- Also see DataONE, Entrez databases, KNB, and PANGAEA, listed under Multidisciplinary repositories.
- The Arabidopsis Information Resource - The Arabidopsis Information Resource (TAIR) maintains a database of genetic and molecular biology datafor the model higher plant Arabidopsis thaliana.
- Array Express(perma.cc). Archive of Functional Genomics Data stores data from high-throughput functional genomics experiments, and provides these data for reuse to the research community.
- Biological Magnetic Resonance Databank(perma.cc). A Repository for Data from NMR Spectroscopy on Proteins, Peptides, Nucleic Acids, and other Biomolecules.
- Biological General Repository for Interaction Datasets - BioGRID. (perma.cc). BioGRID is an interaction repository with data compiled through comprehensive curation efforts.
- BioModels(perma.cc). BioModels is a repository of mathematical models of biological and biomedical systems.
- BOND (Biomolecular Object Network Databank). From Unleashed Informatics.
- Cancer Imaging Archive(perma.cc). TCIA is a service which de-identifies and hosts a large archive of medical images of cancer accessible for public download.
- The Cell: An Image Library Images of all cell types from all organisms, including intracellular structures and movies or animations demonstrating functions. This project relies upon the cell biology community to populate the library. The Cell: An Image Library™ is a freely accessible, easy-to-search, public repository of reviewed and annotated images, videos, and animations of cells from a variety of organisms, showcasing cell architecture, intracellular functionalities, and both normal and abnormal processes. The purpose of this database is to advance research, education, and training, with the ultimate goal of improving human health.
- Coherent X-ray Imaging Data Bank (CXIDB)(perma.cc). The main goal of the Coherent X-ray Imaging Data Bank is to address these problems by creating an open repository for CXI experimental data.
- Crystallography Open Database (COD)(perma.cc). Open-access collection of crystal structures of organic, inorganic, metal-organics compounds and minerals, excluding biopolymers.
- Database of Virulence Factors in Fungal Pathogenes (DFVF)(perma.cc). The database is expected to greatly stimulate and facilitate further studies in fungal pathogens; both experimental biologists and computational biologists can use the database and/or the predicted virulence factors to guide their search for new virulence factors and/or discovery of new pathogen-host interaction mechanisms in fungi.
- Databases at EBI. From the European Bioinformatics Institute (EBI). This is a web directory of the EBI databases. Also see the FTP interface.
- DataBasin. OA data in conservation. From the Conservation Biology Institute in partnership with Rhiza Labs.
- DNA Databank of Japan - DDBJ(perma.cc). Bioinformation and DDBJ Center provides sharing and analysis services for data from life science researches and advances science.
- dbGaP(perma.cc). The database of Genotypes and Phenotypes (dbGaP) was developed to archive and distribute the data and results from studies that have investigated the interaction of genotype and phenotype in Humans.
- dpSNP(perma.cc). dbSNP contains human single nucleotide variations, microsatellites, and small-scale insertions and deletions along with publication, population frequency, molecular consequence, and genomic and RefSeq mapping information for both common variations and clinical mutations.
- dbVar(perma.cc). dbVar is NCBI's database of human genomic Structural Variation — large variants >50 bp including insertions, deletions, duplications, inversions, mobile elements, translocations, and complex variants.
- Dryad Dryad is an international repository of data underlying scientific and medical publications, particularly data for which no specialized repository exists. All material in Dryad is associated with a scholarly publication. Most data in the repository are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted. Dryad is a non-profit organization.
- Electron Microscopy Data Bank (EMDB)(perma.cc). Global resource for 3-Dimensional Electron Microscopy (3DEM) structure data archiving and retrieval, news, events, software tools, data standards, validation methods, and community challenges.
- Eukaryotic Pathogen Database Resources (EuPathDB)(perma.cc). EuPathDB (formerly ApiDB) is an integrated database covering the eukaryotic pathogens in the genera listed in the [EuPathDB (formerly ApiDB) is an integrated database covering the eukaryotic pathogens in the genera listed in our Data Summary page Data Summary] page.
- The European Genome-phenome Archive (EGA)(perma.cc). The European Genome-phenome Archive (EGA) is a service for permanent archiving and sharing of all types of personally identifiable genetic and phenotypic data resulting from biomedical research projects.
- European Variation Archive(perma.cc). An open-access database of all types of genetic variation data from all species.
- G-Node GIN(perma.cc). Modern Research Data Management for Neuroscience.
- Gene Expression Omnibus High-throughput functional genomic data, including all array-based applications and some high-throughput sequencing data.
- Global Biodiversity Information Facility (GBIF) (perma.cc). "Free and open access to biodiversity data." Data portal launched in 2007 by institutions in 17 countries under a non-binding inter-governmental agreement.
- Human Protein Atlas(perma.cc). All the data in the knowledge resource is open access to allow scientists both in academia and industry to freely access the data for exploration of the human proteome.
- http://idr.openmicroscopy.org/about/ Image Data Repository - IDR(perma.cc). The Image Data Resource (IDR) is a public repository of reference image datasets from published scientific studies. IDR enables access, search and analysis of these highly annotated datasets.
- ImmPort Shared Data(perma.cc). The ImmPort project provides advanced information technology support in the archiving and exchange of scientific data for the diverse community of life science researchers supported by NIAID/DAIT and serves as a long-term, sustainable archive of research and clinical data.
- Influenza Research Database(perma.cc). This resource contains avian and non-human mammalian influenza surveillance data, human clinical data associated with virus extracts, phenotypic characteristics of viruses isolated from extracts, and all genomic and proteomic data available in public repositories for influenza viruses.
- Integrated Taxonomic Information System ITIS(perma.cc). It provides authoritative taxonomic information on plants, animals, fungi, and microbes of North America and the world.
- MetaboLights(perma.cc). MetaboLights is a database for Metabolomics experiments and derived information.
- MGnify(perma.cc). MGnify offers an automated pipeline for the analysis and archiving of microbiome data to help determine the taxonomic diversity and functional & metabolic potential of environmental samples.
- Molecular Biology Databases. From Shirley Fung. A list of 34 databases with annotations to show their openness under six criteria. Also see her list of 7 databases which comply with the Science Commons Open Access Data Protocol.
- MorphoBank(perma.cc). "Homology of phenotypes over the web." Hosted by the State University of New York at Stony Brook. MorphoBank assists scientists building the Tree of Life - the genealogy of all living and extinct species.
- Mouse Genome Informatics (MGI)(perma.cc). MGI is the international database resource for the laboratory mouse, providing integrated genetic, genomic, and biological data to facilitate the study of human health and disease.
- Movebank Data Repository(perma.cc). Movebank is a free, online database of animal tracking data hosted by the Max Planck Institute of Animal Behavior.
- National Biological Information Infrastructure A broad, collaborative program to provide increased access to data and information on the nation's biological resources. The NBII links diverse, high-quality biological databases, information products, and analytical tools maintained by NBII partners and other contributors in government agencies, academic institutions, non-government organizations, and private industry. (Note: In the President's budget for Fiscal Year 2012 the repository was terminated.)
- NCBI Taxonomy(perma.cc). The Taxonomy Database is a curated classification and nomenclature for all of the organisms in the public sequence databases.
- The Network Data Exchange (NDEx)(perma.cc). The NDEx Project provides an open-source framework where scientists and organizations can share, store, manipulate, and publish biological network knowledge.
- NeuroMorpho.org(perma.cc). NeuroMorpho.Org is a centrally curated inventory of digitally reconstructed neurons associated with peer-reviewed publications.
- PaleoBiology Database. "We are bringing together taxonomic and distributional information about the entire fossil record of plants and animals." From a large number of researchers at a large number of institutions.
- PeptideAtlas(perma.cc). A multi-organism, publicly accessible compendium of peptides identified in a large set of tandem mass spectrometry proteomics experiments.
- Peptidome. For "tandem mass spectrometry peptide and protein identification data." From the US National Center for Biotechnology Information.
- Planet A network of European Plant Databases.
- PRIDE Proteomics IDentifications Database(perma.cc). This service is part of the ELIXIR infrastructure.
- Protein Circular Dichroism Data Bank (PCDDB)(perma.cc). The Protein Circular Dichroism Data Bank (PCDDB) is a public repository that archives and freely distributes circular dichroism (CD) and synchrotron radiation CD (SRCD) spectral data and their associated experimental metadata.
- RCSB Protein Data Bank. From the Research Collaboratory for Structural Bioinformatics (RCSB).
- ProteomeXChange(perma.cc). The ProteomeXchange Consortium was established to provide globally coordinated standard data submission and dissemination pipelines involving the main proteomics repositories, and to encourage open data policies in the field.
- Rat Genome Database(perma.cc). The Rat Genome Database (RGD) was established in 1999.
- SABIO Biochemical Reaction Kinetics Database(perma.cc). SABIO-RK is a curated database that contains information about biochemical reactions, their kinetic rate equations with parameters and experimental conditions.
- Small Angle Scattering Biological Data Bank(perma.cc). Curated repository for small angle scattering data and models.
- Structural Biology Data Grid. (perma.cc). It supports publication of X-ray diffraction, MicroED, LLSM datasets, as well as structural models.
- TreeBASE. "A Database of Phylogenetic Knowledge." Released in March 2010 based on a prototype launched in 1994. Hosted by the Phyloinformatics Research Foundation.
- The Universal Protein Resource (UniProt) is a comprehensive resource for protein sequence and annotation data. The UniProt databases are the UniProt Knowledgebase (UniProtKB), the UniProt Reference Clusters (UniRef), and the UniProt Archive (UniParc). The UniProt Metagenomic and Environmental Sequences (UniMES) database is a repository specifically developed for metagenomic and environmental data.
- VectorBase(perma.cc). A National Institute of Allergy and Infectious Diseases (NIAID) Bioinformatics Resource Center (BRC) providing genomic, phenotypic and population-centric data to the scientific community for invertebrate vectors of human pathogens.
- Worldwide Protein Data Bank (wwPDB)(perma.cc). The Protein Data Bank archive (PDB) has served as the single repository of information about the 3D structures of proteins, nucleic acids, and complex assemblies.
- Zebrafish Model Organism Database (ZFIN)(perma.cc). A database of genetic and genomic data for the zebrafish (Danio rerio) as a model organism.
Chemistry
- Also see BCO-DMO, Marine Biology data, listed with Marine Sciences repositories.
- Also see Entrez databases, listed under Multidisciplinary repositories.
- Cambridge Structural Database The CCDC is a non-profit, charitable Institution whose objectives are the general advancement and promotion of the science of chemistry and crystallography for the public benefit.
- ChemSpider. Hosted by the Royal Society of Chemistry.
- ChemStar. Maintained by India's National Chemical Laboratory and sponsored by India's Department for Scientific & Industrial Research.
- ChemSynthesis. A database of chemicals and their physical properties.
- ChemXSeer. Hosted by Pennsylvania State University.
- Cooperative Association for Internet Data Analysis (CAIDA) Archive of data for scientific analysis of network functions.
- CrystalEye. From the Unilever Cambridge Centre for Molecular Informatics at the University of Cambridge.
- Crystallography Open Database. A joint project of the Mineralogical Society of America, Mineralogical Association of Canada, European Journal of Mineralogy, International Union of Crystallography, and the US National Science Foundation. Data are in the public domain.
- eCrystals. From the Southampton Chemical Crystallography Group and the EPSRC UK National Crystallography Service.
- NMRShiftDB. For "organic structures and their nuclear magnetic resonance (nmr) spectra." Distributed nodes from the EBI, University of Mainz and the Max Plank Institute for Chemical Ecology. Data are license with the GNU FDL.
- ChemStar. Maintained by India's National Chemical Laboratory and sponsored by India's Department for Scientific & Industrial Research.
- ioChem-BD Computational Chemistry Datasets(perma.cc). The Computational Chemistry Results Repository.
- Open Notebook Science Solubility Challenge. Maintained by Jean-Claude Bradley, Rajarshi Guha, Andrew Lang and Cameron Neylon. A database of non-aqueous solubility measurements with links to lab notebook pages where experiments were recorded. The database can be searched via Web Query or alternate means.
- PubChem. From the U.S. National Center for Biotechnology Information of the National Institutes of Health (NIH).
- WorldWideMolecularMatrix. "An Open collection of information on small molecules." From the University of Cambridge.
- ZINC. "A free database of commercially-available compounds for virtual screening." From the Shoichet Laboratory in the Department of Pharmaceutical Chemistry at the University of California, San Francisco.
Computer Science
- CiteSeerX provides its databases of nearly 2 million documents and the associated texts and pdfs for research.
- Cooperative Association for Internet Data Analysis (CAIDA) Archive of data for scientific analysis of network functions.
- FreeStatistics of Irreproducible Research(perma.cc). The purpose of this project is to facilitate the creation, maintenance, and permanent storage of statistical computation objects that empower authors to publish reproducible and reusable research (in the form of a Compendium) through a series of web services.
- GitHub keeps your public and private code available, secure, and backed up.
- Google Code Project Hosting Project Hosting on Google Code provides a free collaborative development environment for open source projects. Each project comes with its own member controls, Subversion/Mercurial repository, issue tracker, wiki pages, and downloads section. Our project hosting service is simple, fast, reliable, and scalable, so that you can focus on your own open source development.
- Launchpad can host your project’s source code using the Bazaar version control system. We also import over 2000 CVS, SVN, Git and Mercurial projects, so you can use Bazaar with those too.
- ReproZip!(perma.cc). ReproZip can automatically pack your research along with all necessary data files, libraries, environment variables and options into a self-contained bundle. Then ReproZip can use that bundle to automatically set up the same original environment so anybody can reproduce the research on a different machine, without tracking down and installing the dependencies, or even having to run the same operating system.
- SourceForge 2.7 million developers create powerful software in over 260,000 projects. Our popular directory connects more than 46 million consumers with these open source projects and serves more than 2,000,000 downloads a day. SourceForge is where open source happens.
- SNAP Stanford Large Network Dataset Collection. The SNAP library is being actively developed since 2004 and is organically growing as a result of our research pursuits in analysis of large social and information networks. Largest network we analyzed so far using the library was the Microsoft Instant Messenger network from 2006 with 240 million nodes and 1.3 billion edges.
- KONECT (the Koblenz Network Collection) is a project to collect large network datasets of all types in order to perform research in network science and related fields, collected by the Institute of Web Science and Technologies at the University of Koblenz–Landau.
- pajek's network data sources.
Energy
- DOE Data Explorer. From the US Department of Energy (DOE). Data generated by DOE-sponsored research.
- Energy Harvesting Network Data Repository(perma.cc). An EPSRC Funded Network.
- OpenEI: Open Energy Information. Freely-available energy data, tools, models, and other resources.
Engineering
- Also see Multidisciplinary repositories.
- TRID(perma.cc). an integrated database that combines the records from TRB’s Transportation Research Information Services (TRIS) Database and the OECD’s Joint Transport Research Centre’s International Transport Research Documentation (ITRD) Database. TRID provides access to more than 1.25 million records of transportation research worldwide.
Environmental sciences
- Also see BCO-DMO, Marine Biology data, listed with Marine Sciences repositories.
- Also see DataONE, KNB, and PANGAEA, listed under Multidisciplinary repositories.
- Also see Dryad, listed with Biology repositories.
- British Atmospheric Data Centre (BADC). From the Natural Environment Research Council (NERC). Many datasets are openly accessible but some are restricted.
- Climate Change Data Portal. From the Environment Department of the World Bank.
- Climate Data. A section within the Comprehensive Knowledge Archive Network of the Open Knowledge Foundation (OKF). A joint project of the OKF, Climate Code and Real Climate.
- California Water CyberInfrastructure. Hydrology data on California's watersheds. From the Berkeley Water Center.
- Consortium of Universities for the Advancement of Hydrologic Science, Inc HIS stands for Hydrologic Information System. CUAHSI's HIS is an internet based system to support the sharing of hydrologic data. It consists of databases connected using the internet through web services as well as software for data discovery, access and publication.
- EDI Data Portal(perma.cc). The EDI Data Portal contains environmental and ecological data packages contributed by a number of participating organizations.
- The Marine Geoscience Data System (MGDS)(perma.cc). The Marine Geoscience Data System (MGDS) provides access to data portals for the NSF-supported Ridge 2000 and MARGINS programs, the Antarctic and Southern Ocean Data Synthesis, the Global Multi-Resolution Topography Synthesis, and Seismic Reflection Field Data Portal.
- National Snow and Ice Data Center (NSIDC) Cryospheric datasets from ground field research and satellites.
- National Ecological Observatory Network (NEON). A joint project of 50+ US universities and laboratories.
- NERC Data Centers(perma.cc). NERC has a network of environmental data centres that provide a focal point for NERC's scientific data and information. These centres hold data from environmental scientists working in the UK and around the world.
- Polar Data Catalogue A primarily Canadian archive of free RADARSAT imagery as well as Arctic, Antarctic, and other cryospheric data sets covering a range of disciplines, from natural sciences and policy to health and social sciences.
- Socioeconomic Data and Applications Center (SEDAC) specializes in spatial data and services in support of human-environment research and applications, in the context of NASA’s Earth science mission and the overall U.S. Global Change Research Program.
Geology
- Also see PANGAEA, listed under Multidisciplinary repositories.
- GSA Data Repository. From the Geological Society of America.
- IRIS (Incorporated Research Institutions for Seismology). From 100+ US universities and the National Science Foundation.
Geosciences and geospatial data
- Also see DataONE and PANGAEA, listed under Multidisciplinary repositories.
- Commons of Geographic Data. "This site is intended for any data in any format that can be referenced to location on the earth." From the University of Maine.
- EarthChem Library(perma.cc). The EarthChem Library is a data repository that archives, publishes and makes accessible data and other digital content from geoscience research (analytical data, data syntheses, models, technical reports, etc).
- GeoCommons. From FortiusOne.
- GeoNames. A database of placenames, under a CC-BY license. Founded by Marc Wick.
- The Geosciences Network (GEON) project is a collaboration among a dozen PI institutions and a number of other partner projects, institutions, and agencies to develop cyberinfrastructure in support of an environment for integrative geoscience research. GEON is funded by the NSF Information Technology Research (ITR) program.
- Magnetics Information Consortium (MagIC)(perma.cc). Improves research capacity in the Earth and Ocean sciences by maintaining an open community digital data archive for rock and paleomagnetic data with portals that allow users access to archive, search, visualize, download, and combine these versioned datasets.
- National Geographic Data Center Archive of national and international marine environmental and ecosystem datasets.
- National Space Science Data Center serves as the permanent archive for NASA space science mission data. "Space science" means astronomy and astrophysics, solar and space plasma physics, and planetary and lunar science. As permanent archive, NSSDC teams with NASA's discipline-specific space science "active archives" which provide access to data to researchers and, in some cases, to the general public.
- OpenTopography(perma.cc). OpenTopography facilitates community access to high-resolution, Earth science-oriented, topography data, and related tools and resources.
- Polar Data Catalogue A primarily Canadian archive of free RADARSAT imagery as well as Arctic, Antarctic, and other cryospheric data sets covering a range of disciplines, from natural sciences and policy to health and social sciences.
- ShareGeo. Integrating the older GRADE (Geospatial Repository for Academic Deposit and Extraction) repository. From EDINA. (Repository discontinued.)
Linguistics
- See the 40+ members of the Open Language Archives Community (OLAC).
- TROLLing. Hosted by UiT. TROLLing "is designed as an archive of linguistic data and statistical code. The archive is open access, which means that all information is available to to everyone. All postings are accompanied by searchable metadata that identify the researchers, the languages and linguistic phenomena involved, the statistical methods applied, and scholarly publications based on the data (where relevant). Linguists worldwide are invited to post datasets and statistical models used in linguistic research."
Marine sciences
- Also see DataONE and PANGAEA, listed under Multidisciplinary repositories.
- BCO-DMO. The Biological and Chemical Oceanography Data Management Office, provides access to data sets contributed by investigators funded by the Biological and Chemical Oceanography sections of the US National Science Foundation (NSF).
- Naval Oceanography Portal Data Services. From the United States Naval Observatory (USNO).
- SeaDataNet. Funded by the EU and coordinated by Institut Français de Recherche pour l'Exploitation de la Mer (IFREMER).
- SEAONE - Sea Open Scientific Data Publication(perma.cc). SEANOE (SEA scieNtific Open data Edition) is a publisher of scientific data in the field of marine sciences. Data published by SEANOE are available free. They can be used in accordance with the terms of the Creative Commons license selected by the author of data.
Medicine
- Also see Entrez databases, listed under Multidisciplinary repositories.
- caNanoLab(perma.cc) A data sharing portal designed to facilitate information sharing across the international biomedical nanotechnology research community to expedite and validate the use of nanotechnology in biomedicine.
- [ebi.ac.uk/chembl/ CheMBL]. ChEMBL is a manually curated database of bioactive molecules with drug-like properties. It brings together chemical, bioactivity and genomic data to aid the translation of genomic information into effective new drugs.
- Dryad Dryad is an international repository of data underlying scientific and medical publications, particularly data for which no specialized repository exists. All material in Dryad is associated with a scholarly publication. Most data in the repository are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted. Dryad is a non-profit organization.
- FlowRepository(perma.cc). FlowRepository is a database of flow cytometry experiments where you can query and download data collected and annotated according to the MIFlowCyt standard.
- GenBank. From the U.S. National Center for Biotechnology Information of the National Institutes of Health.
- Gene Expression Omnibus. From the U.S. National Center for Biotechnology Information of the National Institutes of Health.
- The Health and Medical Care Archive (HMCA) is the data archive of the Robert Wood Johnson Foundation (RWJF), the largest philanthropy devoted exclusively to health and health care in the United States. Operated by the Inter-university Consortium for Political and Social Research (ICPSR) at the University of Michigan, HMCA preserves and disseminates data collected by selected research projects funded by the Foundation and facilitates secondary analyses of the data. The data collections in HMCA include surveys of health care professionals and organizations, investigations of access to medical care, surveys on substance abuse, and evaluations of innovative programs for the delivery of health care. Our goal is to increase understanding of health and health care in the United States through secondary analysis of RWJF-supported data collections.
- MIRAGE (Middlesex medical Image Repository with a CBIR ArchivinG Environment). From JISC and Middlesex University.
- Melanoma Molecular Map Project. On melanoma biology and treatment.
- [National Addiction & HIV Data Archive Program (NAHDAP)](perma.cc). The scope of the data housed at NAHDAP covers a wide range of legal and illicit drugs (alcohol, tobacco, marijuana, cocaine, synthetic drugs, and others) and the trajectories, patterns, and consequences of drug use as well as related predictors and outcomes.
- National Center for Biotechnology Information (NCBI) The National Center for Biotechnology Information advances science and health by providing access to biomedical and genomic information.
- National Database for Autism Research (NDAR)(perma.cc). The National Institute of Mental Health Data Archive (NDA) makes available human subjects data collected from hundreds of research projects across many scientific domains.
- NeuroMorpho. Neuronal morphology data. From the Krasnow Institute for Advanced Study at George Mason University.
- Ocean Tool for Public Understanding and Science(perma.cc). OcToPUS relies on established free and open-source geospatial technology to provide interactive access to dynamically updated, multi-dimensional data on the marine environment.
- OpenTrials. OpenTrials is a repository of clinical trial data hosted by Open Knowledge International.
- Project Data Sphere, LLC, is a repository to broadly share, integrate and analyze historical, de-identified, patient-level data from academic and industry cancer Phase II-III clinical trials. Access to the Project Data Sphere platform is available to researchers affiliated with life science companies, hospitals and institutions, as well as independent researchers, at no cost and without requiring a research proposal.
- SICAS Medical Image Repository(perma.cc). A place to store medical research data.
- Vivli(perma.cc). From the Center for Global Clinical Research Data. The Vivli platform includes an independent data repository, in-depth search engine and a secure research environment.
Multidisciplinary repositories
- Also see Social Sciences.
- Also see BCO-DMO, Marine Biology data, listed with Marine Sciences repositories.
- 3TU.Datacentre. A consortial data repository for Delft University of Technology, Eindhoven University of Technology and the University of Twente.
- Data Archiving and Networked Services. Dutch research data in the humanities and social sciences. From the Royal Netherlands Academy of Arts and Sciences (KNAW) and the Netherlands Organisation for Scientific Research (NWO).
- DataCite (perma.cc). DataCite is a leading global non-profit organisation that provides persistent identifiers (DOIs) for research data and other research outputs.
- Data Conservancy(perma.cc). Data Conservancy is devoted to developing institutional solutions for the challenges of data collection, preservation and re-use.
- DataHub(perma.cc). There are thousands of datasets from financial market data and population growth to cryptocurrency prices.
- DataONE DataONE is an international federation of data repositories containing earth observations data, including data from fields such as ecology, biology, evolution, and environmental sciences such as hydrology, oceanography, and atmospheric science. DataONE is a federation with participation from hundreds of field stations, universities, and government agencies through the DataONE Member Nodes.
- The Dataverse Network. From Harvard's Institute for Quantitative Social Science (IQSS).
- Dryad Dryad is an international repository of data underlying scientific and medical publications, particularly data for which no specialized repository exists. All material in Dryad is associated with a scholarly publication. Most data in the repository are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted. Dryad is a non-profit organization.
- EASY(perma.cc). EASY offers sustainable archiving of research data and access to thousands of datasets.
- Edinburgh DataShare hosted by Edinburgh University Data Library. A repository for data produced by research at the University of Edinburgh.
- Entrez databases. A directory of chemical, biochemical, biomedical, and medical databases from the U.S. National Center for Biotechnology Information of the National Institutes of Health.
- EUDAT(perma.cc). EUDAT offers heterogeneous research data management services and storage resources, supporting multiple research communities as well as individuals, through a geographically distributed, resilient network distributed across 15 European nations and data is stored alongside some of Europe’s most powerful supercomputers.
- FigShare. Scientific publishing as it stands is an inefficient way to do science on a global scale. A lot of time and money is being wasted by groups around the world duplicating research that has already been carried out. FigShare allows you to share all of your data, negative results and unpublished figures. In doing this, other researchers will not duplicate the work, but instead may publish with your previously wasted figures, or offer collaboration opportunities and feedback on preprint figures.
- KNB The Knowledge Network for Biocomplexity (KNB) is an international data repository containing ecology, biology, and environmental science data with a global distribution. The KNB is a grass-roots partnership of collaborating feld stations, laboratories, and research networks that openly publish and share data. Founding partners include the National Center for Ecological Analysis and Synthesis (NCEAS) and the Long-term Ecological Research Network (LTER). The KNB is a Member Node within the DataONE data federation.
- KPBC. Regional academic repository for data in all fields. Poland
- Microsoft Research Open Data(perma.cc). A collection of free datasets from Microsoft Research to advance state-of-the-art research in areas such as natural language processing, computer vision, and domain specific sciences. Download or copy directly to a cloud-based Data Science Virtual Machine for a seamless development experience.
- Open Commons Consortium (OCC). The OCC is a not for profit that manages and operates cloud computing and data commons infrastructure to support scientific, medical, health care and environmental research. OCC members span the globe and include over 30 universities, companies, government agencies and national laboratories.
- Open Science Data Cloud (OSDC). The OSDC is a data science ecosystem in which researchers can house and share their own scientific data, access complementary public datasets, build and share customized virtual machines with whatever tools necessary to analyze their data, and perform the analysis to answer their research questions. It is a one-stop shop for making scientific research faster and easier.
- Open Science Framework (OSF) Open Science Framework serves as a scholarly commons for documentation, files, collaboration, and connecting to services for research outputs.
- PANGAEA. "PANGAEA" stands for "Publishing Network for Geoscientific & Environmental Data". Hosted by the Alfred Wegener Institute for Polar and Marine Research and the University of Bremen's Center for Marine Environmental Sciences. Open to deposits from any scientist. Most datasets are open; some are restricted.
- Public Data Sets on AWS. From Amazon Web Services. The site already hosts OA datasets in biology, chemistry, and economics, and is willing to host them in any field.
- Scholars Portal Dataverse. A data repository hosted by Scholars Portal, a consortial service of the Ontario Council of University Libraries in Canada. Open to deposits from any user across all fields of research.
- Science 3.0 Open Data. A repository for RDF datasets in the public domain, in any field. From Science 3.0.
- Scientific Data recommended repositories(perma.cc). Spreadsheet listing data repositories that are recommended by Scientific Data (Springer Nature) as being suitable for hosting data associated with peer-reviewed articles. Please see the repository list on Scientific Data's website for the most up to date list.
- Tromsø Repository of Language and Linguistics (TROLLing)(perma.cc). TROLLing is designed as an archive of linguistic data and statistical code. The archive is open access, which means that all information is available to to everyone. All postings are accompanied by searchable metadata that identify the researchers, the languages and linguistic phenomena involved, the statistical methods applied, and scholarly publications based on the data (where relevant).
- USU Repository University of Sumatera Utara, Medan, Indonesia.
- UPSpace University of Pretoria Research Repository, South Africa.
- Webscope(perma.cc). The Yahoo Webscope Program is a reference library of interesting and scientifically useful datasets for non-commercial use by academics and other scientists. All datasets have been reviewed to conform to Yahoo's data protection standards, including strict controls on privacy.
Physics
- Also see Astronomy.
- Blue Obelisk Data Repository. Repository of isotope masses, under MIT license. From the Blue Obelisk. Described in 10.1021/ci050400b.
- CERN Scientific Information Online particle physics data and information
- HEP Data The data comprise total and differential cross sections, structure functions, fragmentation functions, distribuitions of jet measures, polarisations, etc... from a wide range of interactions.
- LAMBDA(perma.cc). LAMBDA is a part of NASA's High Energy Astrophysics Science Archive Research Center (HEASARC).
- Nist Atomic Spectra Database The Atomic Spectra Database (ASD) contains data for radiative transitions and energy levels in atoms and atomic ions. Data are included for observed transitions of 99 elements and energy levels of 56 elements.
Social sciences
- Also see Multidisciplinary repositories.
- Archeology Data Service(perma.cc). Heritage data, with over 20 years of experience supporting research, learning and teaching with free, high quality and dependable digital resources.
- Association of Religion Data Archives Coverage includes international surveys, U.S. church membership data, and U.S. Surveys.
- Australian Social Science Data Archive. From the Australian Demographic and Social Research Institute at the Australian National University.
- CESSDA Data Portal. From the Council of European Social Science Data Archives (CESSDA).
- Databrary A repository for sharing and reusing research video data and related metadata in the developmental and learning sciences. Hosted at New York University with support from The Pennsylvania State University.
- Digital Repositories E-Science Network (DReSNeT). From the UK Engineering & Physical Sciences Research Council (EPSRC). A network of social science repositories for texts and data.
- Economic and Social Science Data Service. From the UK Data Archive (UKDA) and Institute for Social and Economic Research (ISER), University of Essex; Manchester Information and Associated Services (MIMAS), and the Cathie Marsh Centre for Census and Survey Research (CCSR), University of Manchester. Access to data requires registration.
- European Nucleotide Archive(perma.cc). The European Nucleotide Archive (ENA) provides a comprehensive record of the world's nucleotide sequencing information, covering raw sequencing data, sequence assembly information and functional annotation.
- ICPSR (Inter-University Consortium for Political and Social Research). At the University of Michigan.
- National Archive of Criminal Justice Data holds over 700 data collections relating to criminal justice.
- NOMAD Repository(perma.cc). Host, organize, and share materials data.
- openICPRS(perma.cc). openICPSR is a great place to share and store your social and behavioral science research data. Your data will be preserved as-is and be available to data users at no cost.
- Qualitative Data Repository(perma.cc). QDR curates, stores, preserves, publishes, and enables the download of digital data generated through qualitative and multi-method research in the social sciences.
- Roper Center for Public Opinion Research data from surveys of public opinion from the 1930s to the present.