Difference between revisions of "Data repositories"

From Open Access Directory
Jump to navigation Jump to search
(38 intermediate revisions by the same user not shown)
Line 42: Line 42:
  
 
* [http://nssdc.gsfc.nasa.gov/ National Space Science Data Center].  From the US [http://www.nasa.gov/ National Aeronautics and Space Administration] (NASA).
 
* [http://nssdc.gsfc.nasa.gov/ National Space Science Data Center].  From the US [http://www.nasa.gov/ National Aeronautics and Space Administration] (NASA).
 +
 +
* [http://starchive.org/ Starchive]([https://perma.cc/7AEH-WA8U perma.cc]). An open source, open access stelar archive.
  
 
== Biology ==
 
== Biology ==
Line 49: Line 51:
  
 
* [http://www.arabidopsis.org/submit/index.jsp The Arabidopsis Information Resource] - The Arabidopsis Information Resource (TAIR) maintains a [http://www.arabidopsis.org/search/ERwin/Tair.htm database] of genetic and [http://www.arabidopsis.org/about/datasources.jsp molecular biology data]for the model higher plant [http://www.arabidopsis.org/portals/education/aboutarabidopsis.jsp ''Arabidopsis thaliana''].
 
* [http://www.arabidopsis.org/submit/index.jsp The Arabidopsis Information Resource] - The Arabidopsis Information Resource (TAIR) maintains a [http://www.arabidopsis.org/search/ERwin/Tair.htm database] of genetic and [http://www.arabidopsis.org/about/datasources.jsp molecular biology data]for the model higher plant [http://www.arabidopsis.org/portals/education/aboutarabidopsis.jsp ''Arabidopsis thaliana''].
 +
 +
* [https://thebiogrid.org/ Biological General Repository for Interaction Datasets] - BioGRID. ([https://perma.cc/34CE-XK79 perma.cc]). BioGRID is an interaction repository with data compiled through comprehensive curation efforts.
  
 
* [http://bond.unleashedinformatics.com/ BOND] (Biomolecular Object Network Databank). From [http://www.unleashedinformatics.com/ Unleashed Informatics].
 
* [http://bond.unleashedinformatics.com/ BOND] (Biomolecular Object Network Databank). From [http://www.unleashedinformatics.com/ Unleashed Informatics].
  
 
* [http://cellimagelibrary.org/pages/contribute The Cell: An Image Library] Images of all cell types from all organisms, including intracellular structures and movies or animations demonstrating functions. This project relies upon the cell biology community to populate the library. The Cell: An Image Library™ is a freely accessible, easy-to-search, public repository of reviewed and annotated images, videos, and animations of cells from a variety of organisms, showcasing cell architecture, intracellular functionalities, and both normal and abnormal processes. The purpose of this database is to advance research, education, and training, with the ultimate goal of improving human health.
 
* [http://cellimagelibrary.org/pages/contribute The Cell: An Image Library] Images of all cell types from all organisms, including intracellular structures and movies or animations demonstrating functions. This project relies upon the cell biology community to populate the library. The Cell: An Image Library™ is a freely accessible, easy-to-search, public repository of reviewed and annotated images, videos, and animations of cells from a variety of organisms, showcasing cell architecture, intracellular functionalities, and both normal and abnormal processes. The purpose of this database is to advance research, education, and training, with the ultimate goal of improving human health.
 +
 +
* [http://sysbio.unl.edu/DFVF/ Database of Virulence Factors in Fungal Pathogenes] (DFVF)([https://perma.cc/NK4J-Z73P perma.cc]). The database is expected to greatly stimulate and facilitate further studies in fungal pathogens; both experimental biologists and computational biologists can use the database and/or the predicted virulence factors to guide their search for new virulence factors and/or discovery of new pathogen-host interaction mechanisms in fungi.
  
 
* [http://www.ebi.ac.uk/Information/databases_sitemap.html Databases at EBI].  From the [http://www.ebi.ac.uk/ European Bioinformatics Institute] (EBI).  This is a web directory of the EBI databases.  Also see the [ftp://ftp.ebi.ac.uk/pub/databases/ FTP interface].
 
* [http://www.ebi.ac.uk/Information/databases_sitemap.html Databases at EBI].  From the [http://www.ebi.ac.uk/ European Bioinformatics Institute] (EBI).  This is a web directory of the EBI databases.  Also see the [ftp://ftp.ebi.ac.uk/pub/databases/ FTP interface].
  
 
* [http://databasin.org/ DataBasin].  OA data in conservation.  From the [http://www.consbio.org/ Conservation Biology Institute] in partnership with [http://www.rhizalabs.com/ Rhiza Labs].
 
* [http://databasin.org/ DataBasin].  OA data in conservation.  From the [http://www.consbio.org/ Conservation Biology Institute] in partnership with [http://www.rhizalabs.com/ Rhiza Labs].
 +
 +
* [https://www.ddbj.nig.ac.jp/index-e.html DNA Databank of Japan - DDBJ]([https://perma.cc/3ZKD-G9RX perma.cc]). Bioinformation and DDBJ Center provides sharing and analysis services for data from life science researches and advances science.
 +
 +
* [https://www.ncbi.nlm.nih.gov/snp/ dpSNP]([https://perma.cc/L3AW-6EXP perma.cc]). dbSNP contains human single nucleotide variations, microsatellites, and small-scale insertions and deletions along with publication, population frequency, molecular consequence, and genomic and RefSeq mapping information for both common variations and clinical mutations.
 +
 +
* [https://www.ncbi.nlm.nih.gov/dbvar/ dbVar]([https://perma.cc/JDA9-K62V perma.cc]). dbVar is NCBI's database of human genomic Structural Variation — large variants >50 bp including insertions, deletions, duplications, inversions, mobile elements, translocations, and complex variants.
  
 
* [http://www.datadryad.org/ Dryad] Dryad is an international repository of data underlying scientific and medical publications, particularly data for which no specialized repository exists. All material in Dryad is associated with a scholarly publication. Most data in the repository are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted.  Dryad is a non-profit organization.  
 
* [http://www.datadryad.org/ Dryad] Dryad is an international repository of data underlying scientific and medical publications, particularly data for which no specialized repository exists. All material in Dryad is associated with a scholarly publication. Most data in the repository are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted.  Dryad is a non-profit organization.  
 +
 +
* [https://www.ebi.ac.uk/eva/ European Variation Archive]([https://perma.cc/NH45-7GPB perma.cc]). An open-access database of all types of genetic variation data from all species.
  
 
* [http://www.ncbi.nlm.nih.gov/geo/info/submission.html Gene Expression Omnibus] High-throughput functional genomic data, including all array-based applications and some high-throughput sequencing data.
 
* [http://www.ncbi.nlm.nih.gov/geo/info/submission.html Gene Expression Omnibus] High-throughput functional genomic data, including all array-based applications and some high-throughput sequencing data.
  
* [http://www.gbif.org/ Global Biodiversity Information Facility] (GBIF).  "Free and open access to biodiversity data."  Data portal launched in 2007 by institutions in 17 countries under a non-binding inter-governmental agreement.
+
* [http://www.gbif.org/ Global Biodiversity Information Facility] (GBIF) ([https://perma.cc/BGE5-CBFH perma.cc]).  "Free and open access to biodiversity data."  Data portal launched in 2007 by institutions in 17 countries under a non-binding inter-governmental agreement.
 +
 
 +
* [https://www.proteinatlas.org/ Human Protein Atlas]([https://perma.cc/2FML-ANKD perma.cc]). All the data in the knowledge resource is open access to allow scientists both in academia and industry to freely access the data for exploration of the human proteome.
 +
 
 +
* [https://www.ebi.ac.uk/metagenomics/ MGnify]([https://perma.cc/FDA2-58XY perma.cc]). MGnify offers an automated pipeline for the analysis and archiving of microbiome data to help determine the taxonomic diversity and functional & metabolic potential of environmental samples.
  
 
* [http://shirleyfung.com/mbdb/ Molecular Biology Databases].  From Shirley Fung. A list of [http://shirleyfung.com/mbdb/filter.php?by=alltab 34 databases] with annotations to show their openness under six criteria.  Also see her list of [http://shirleyfung.com/mbdb/filter.php?by=compliant 7 databases] which comply with the [http://sciencecommons.org/projects/publishing/open-access-data-protocol/ Science Commons Open Access Data Protocol].
 
* [http://shirleyfung.com/mbdb/ Molecular Biology Databases].  From Shirley Fung. A list of [http://shirleyfung.com/mbdb/filter.php?by=alltab 34 databases] with annotations to show their openness under six criteria.  Also see her list of [http://shirleyfung.com/mbdb/filter.php?by=compliant 7 databases] which comply with the [http://sciencecommons.org/projects/publishing/open-access-data-protocol/ Science Commons Open Access Data Protocol].
Line 77: Line 95:
  
 
* [http://www.rcsb.org/pdb/home/home.do RCSB Protein Data Bank].  From the [http://home.rcsb.org/ Research Collaboratory for Structural Bioinformatics] (RCSB).
 
* [http://www.rcsb.org/pdb/home/home.do RCSB Protein Data Bank].  From the [http://home.rcsb.org/ Research Collaboratory for Structural Bioinformatics] (RCSB).
 +
 +
* [http://sabio.h-its.org/ SABIO Biochemical Reaction Kinetics Database]([https://perma.cc/5X9B-H9G2 perma.cc]). SABIO-RK is a curated database that contains information about biochemical reactions, their kinetic rate equations with parameters and experimental conditions.
 +
 +
* [https://www.sasbdb.org/ Small Angle Scattering Biological Data Bank]([https://perma.cc/2248-7JQT perma.cc]). Curated repository for small angle scattering data and models.
  
 
* [http://www.treebase.org TreeBASE]. "A Database of Phylogenetic Knowledge."  Released in March 2010 based on a prototype launched in 1994.  Hosted by the [http://www.phylorf.org/ Phyloinformatics Research Foundation].
 
* [http://www.treebase.org TreeBASE]. "A Database of Phylogenetic Knowledge."  Released in March 2010 based on a prototype launched in 1994.  Hosted by the [http://www.phylorf.org/ Phyloinformatics Research Foundation].
  
*[http://www.uniprot.org The Universal Protein Resource (UniProt)] is a comprehensive resource for protein sequence and annotation data. The UniProt databases are the UniProt Knowledgebase (UniProtKB), the UniProt Reference Clusters (UniRef), and the UniProt Archive (UniParc). The UniProt Metagenomic and Environmental Sequences (UniMES) database is a repository specifically developed for metagenomic and environmental data..
+
*[http://www.uniprot.org The Universal Protein Resource (UniProt)] is a comprehensive resource for protein sequence and annotation data. The UniProt databases are the UniProt Knowledgebase (UniProtKB), the UniProt Reference Clusters (UniRef), and the UniProt Archive (UniParc). The UniProt Metagenomic and Environmental Sequences (UniMES) database is a repository specifically developed for metagenomic and environmental data.
 +
 
 +
* [http://www.ubio.org/index.php?pagename=home uBio]([https://perma.cc/HNR7-RUMU perma.cc]). uBio uses names and taxonomic intelligence to manage information about organisms.
  
 
== Chemistry ==
 
== Chemistry ==
Line 122: Line 146:
  
 
* [http://www.caida.org/data/ Cooperative Association for Internet Data Analysis (CAIDA)] Archive of data for scientific analysis of network functions.
 
* [http://www.caida.org/data/ Cooperative Association for Internet Data Analysis (CAIDA)] Archive of data for scientific analysis of network functions.
 +
 +
* [https://www.freestatistics.org/ FreeStatistics of Irreproducible Research]([https://perma.cc/2MJ6-69LP perma.cc]). The purpose of this project is to facilitate the creation, maintenance, and permanent storage of statistical computation objects that empower authors to publish reproducible and reusable research (in the form of a Compendium) through a series of web services.
  
 
* [https://github.com GitHub] keeps your public and private code available, secure, and backed up.
 
* [https://github.com GitHub] keeps your public and private code available, secure, and backed up.
Line 128: Line 154:
  
 
* [https://code.launchpad.net Launchpad] can host your project’s source code using the Bazaar version control system. We also import over 2000 CVS, SVN, Git and Mercurial projects, so you can use Bazaar with those too.  
 
* [https://code.launchpad.net Launchpad] can host your project’s source code using the Bazaar version control system. We also import over 2000 CVS, SVN, Git and Mercurial projects, so you can use Bazaar with those too.  
 +
 +
* [https://www.reprozip.org/ ReproZip!]([https://perma.cc/4574-F9M7 perma.cc]). ReproZip can automatically pack your research along with all necessary data files, libraries, environment variables and options into a self-contained bundle. Then ReproZip can use that bundle to automatically set up the same original environment so anybody can reproduce the research on a different machine, without tracking down and installing the dependencies, or even having to run the same operating system.
  
 
* [http://sourceforge.net/ SourceForge] 2.7 million developers create powerful software in over 260,000 projects. Our popular directory connects more than 46 million consumers with these open source projects and serves more than 2,000,000 downloads a day. SourceForge is where open source happens.
 
* [http://sourceforge.net/ SourceForge] 2.7 million developers create powerful software in over 260,000 projects. Our popular directory connects more than 46 million consumers with these open source projects and serves more than 2,000,000 downloads a day. SourceForge is where open source happens.
Line 135: Line 163:
 
* [http://konect.uni-koblenz.de/ KONECT ] (the Koblenz Network Collection) is a project to collect large network datasets of all types in order to perform research in network science and related fields, collected by the Institute of Web Science and Technologies at the University of Koblenz–Landau.
 
* [http://konect.uni-koblenz.de/ KONECT ] (the Koblenz Network Collection) is a project to collect large network datasets of all types in order to perform research in network science and related fields, collected by the Institute of Web Science and Technologies at the University of Koblenz–Landau.
  
* [http://pajek.imfm.si/doku.php?id=data:urls:index pajek's ] network data sources
+
* [http://pajek.imfm.si/doku.php?id=data:urls:index pajek's ] network data sources.
  
 
== Energy ==
 
== Energy ==
  
 
* [http://www.osti.gov/dataexplorer/ DOE Data Explorer].  From the US [http://www.energy.gov/ Department of Energy] (DOE).  Data generated by DOE-sponsored research.
 
* [http://www.osti.gov/dataexplorer/ DOE Data Explorer].  From the US [http://www.energy.gov/ Department of Energy] (DOE).  Data generated by DOE-sponsored research.
 +
 +
* [http://eh-network.org/data/ Energy Harvesting Network Data Repository]([https://perma.cc/TD3Y-772Q perma.cc]). An [https://epsrc.ukri.org/ EPSRC] Funded Network. 
  
 
* [http://en.openei.org/ OpenEI: Open Energy Information].  Freely-available energy data, tools, models, and other resources.
 
* [http://en.openei.org/ OpenEI: Open Energy Information].  Freely-available energy data, tools, models, and other resources.
Line 182: Line 212:
  
 
* [http://geodatacommons.umaine.edu/ Commons of Geographic Data].  "This site is intended for any data in any format that can be referenced to location on the earth."  From the [http://www.umaine.edu/ University of Maine].
 
* [http://geodatacommons.umaine.edu/ Commons of Geographic Data].  "This site is intended for any data in any format that can be referenced to location on the earth."  From the [http://www.umaine.edu/ University of Maine].
 +
 +
* [http://www.earthchem.org/library EarthChem Library]([https://perma.cc/V7TY-3JRU perma.cc]). The EarthChem Library is a data repository that archives, publishes and makes accessible data and other digital content from geoscience research (analytical data, data syntheses, models, technical reports, etc).
  
 
* [http://www.geocommons.com/ GeoCommons].  From [http://www.fortiusone.com/ FortiusOne].
 
* [http://www.geocommons.com/ GeoCommons].  From [http://www.fortiusone.com/ FortiusOne].
Line 195: Line 227:
 
* [http://www.nodc.noaa.gov/General/NODC-Submit/ National Geographic Data Center] Archive of national and international marine environmental and ecosystem datasets.  
 
* [http://www.nodc.noaa.gov/General/NODC-Submit/ National Geographic Data Center] Archive of national and international marine environmental and ecosystem datasets.  
  
* [http://nssdc.gsfc.nasa.gov/nssdc/submitting_data.html The National Space Science Data Center] serves as the permanent archive for NASA space science mission data. "Space science" means astronomy and astrophysics, solar and space plasma physics, and planetary and lunar science. As permanent archive, NSSDC teams with NASA's discipline-specific space science "active archives" which provide access to data to researchers and, in some cases, to the general public.
+
* [http://nssdc.gsfc.nasa.gov/nssdc/submitting_data.html National Space Science Data Center] serves as the permanent archive for NASA space science mission data. "Space science" means astronomy and astrophysics, solar and space plasma physics, and planetary and lunar science. As permanent archive, NSSDC teams with NASA's discipline-specific space science "active archives" which provide access to data to researchers and, in some cases, to the general public.
 +
 
 +
* [https://opentopography.org/ OpenTopography]([https://perma.cc/MXZ3-B9ZE perma.cc]). OpenTopography facilitates community access to high-resolution, Earth science-oriented, topography data, and related tools and resources.
  
 
* [http://www.polardata.ca Polar Data Catalogue] A primarily Canadian archive of free RADARSAT imagery as well as Arctic, Antarctic, and other cryospheric data sets covering a range of disciplines, from natural sciences and policy to health and social sciences.
 
* [http://www.polardata.ca Polar Data Catalogue] A primarily Canadian archive of free RADARSAT imagery as well as Arctic, Antarctic, and other cryospheric data sets covering a range of disciplines, from natural sciences and policy to health and social sciences.
Line 215: Line 249:
  
 
* [http://www.seadatanet.org/ SeaDataNet].  Funded by the EU and coordinated by [http://www.ifremer.fr/ Institut Français de Recherche pour l'Exploitation de la Mer] (IFREMER).
 
* [http://www.seadatanet.org/ SeaDataNet].  Funded by the EU and coordinated by [http://www.ifremer.fr/ Institut Français de Recherche pour l'Exploitation de la Mer] (IFREMER).
 +
 +
* [https://www.seanoe.org/ SEAONE - Sea Open Scientific Data Publication]([https://perma.cc/2Y5N-H2UM perma.cc]). SEANOE (SEA scieNtific Open data Edition) is a publisher of scientific data in the field of marine sciences. Data published by SEANOE are available free. They can be used in accordance with the terms of the Creative Commons license selected by the author of data.
  
 
== Medicine ==
 
== Medicine ==
Line 221: Line 257:
  
 
* [http://www.datadryad.org/ Dryad] Dryad is an international repository of data underlying scientific and medical publications, particularly data for which no specialized repository exists. All material in Dryad is associated with a scholarly publication. Most data in the repository are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted.  Dryad is a non-profit organization.  
 
* [http://www.datadryad.org/ Dryad] Dryad is an international repository of data underlying scientific and medical publications, particularly data for which no specialized repository exists. All material in Dryad is associated with a scholarly publication. Most data in the repository are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted.  Dryad is a non-profit organization.  
 +
 +
* [http://flowrepository.org/ FlowRepository]([https://perma.cc/35NS-MCGB perma.cc]). FlowRepository is a database of flow cytometry experiments where you can query and download data collected and annotated according to the MIFlowCyt standard.
  
 
* [http://www.ncbi.nlm.nih.gov/Genbank/index.html GenBank].  From the U.S. [http://www.ncbi.nlm.nih.gov/ National Center for Biotechnology Information] of the [http://www.nih.gov/ National Institutes of Health].
 
* [http://www.ncbi.nlm.nih.gov/Genbank/index.html GenBank].  From the U.S. [http://www.ncbi.nlm.nih.gov/ National Center for Biotechnology Information] of the [http://www.nih.gov/ National Institutes of Health].
  
 
* [http://www.ncbi.nlm.nih.gov/geo/ Gene Expression Omnibus].  From the U.S. [http://www.ncbi.nlm.nih.gov/ National Center for Biotechnology Information] of the [http://www.nih.gov/ National Institutes of Health].
 
* [http://www.ncbi.nlm.nih.gov/geo/ Gene Expression Omnibus].  From the U.S. [http://www.ncbi.nlm.nih.gov/ National Center for Biotechnology Information] of the [http://www.nih.gov/ National Institutes of Health].
 
* [https://opentrials.net/ OpenTrials]. OpenTrials is a repository of clinical trial data hosted by [https://okfn.org/ Open Knowledge International].
 
  
 
* [http://www.icpsr.umich.edu/icpsrweb/HMCA/index.jsp The Health and Medical Care Archive (HMCA)] is the data archive of the Robert Wood Johnson Foundation (RWJF), the largest philanthropy devoted exclusively to health and health care in the United States. Operated by the Inter-university Consortium for Political and Social Research (ICPSR) at the University of Michigan, HMCA preserves and disseminates data collected by selected research projects funded by the Foundation and facilitates secondary analyses of the data. The data collections in HMCA include surveys of health care professionals and organizations, investigations of access to medical care, surveys on substance abuse, and evaluations of innovative programs for the delivery of health care. Our goal is to increase understanding of health and health care in the United States through secondary analysis of RWJF-supported data collections.  
 
* [http://www.icpsr.umich.edu/icpsrweb/HMCA/index.jsp The Health and Medical Care Archive (HMCA)] is the data archive of the Robert Wood Johnson Foundation (RWJF), the largest philanthropy devoted exclusively to health and health care in the United States. Operated by the Inter-university Consortium for Political and Social Research (ICPSR) at the University of Michigan, HMCA preserves and disseminates data collected by selected research projects funded by the Foundation and facilitates secondary analyses of the data. The data collections in HMCA include surveys of health care professionals and organizations, investigations of access to medical care, surveys on substance abuse, and evaluations of innovative programs for the delivery of health care. Our goal is to increase understanding of health and health care in the United States through secondary analysis of RWJF-supported data collections.  
Line 239: Line 275:
  
 
* [http://neuromorpho.org/neuroMorpho/index.jsp NeuroMorpho].  Neuronal morphology data.  From the [http://krasnow.gmu.edu/ Krasnow Institute for Advanced Study] at [http://www.gmu.edu/ George Mason University].
 
* [http://neuromorpho.org/neuroMorpho/index.jsp NeuroMorpho].  Neuronal morphology data.  From the [http://krasnow.gmu.edu/ Krasnow Institute for Advanced Study] at [http://www.gmu.edu/ George Mason University].
 +
 +
* [https://octopus.zoo.ox.ac.uk/beta/ Ocean Tool for Public Understanding and Science]([https://perma.cc/Z59D-C2YS perma.cc]). OcToPUS relies on established free and open-source geospatial technology to provide interactive access to dynamically updated, multi-dimensional data on the marine environment.
 +
 +
* [https://opentrials.net/ OpenTrials]. OpenTrials is a repository of clinical trial data hosted by [https://okfn.org/ Open Knowledge International].
  
 
* [https://www.projectdatasphere.org Project Data Sphere, LLC,] is a repository to broadly share, integrate and analyze historical, de-identified, patient-level data from academic and industry cancer Phase II-III clinical trials.  Access to the Project Data Sphere platform is available to researchers affiliated with life science companies, hospitals and institutions, as well as independent researchers, at no cost and without requiring a research proposal.
 
* [https://www.projectdatasphere.org Project Data Sphere, LLC,] is a repository to broadly share, integrate and analyze historical, de-identified, patient-level data from academic and industry cancer Phase II-III clinical trials.  Access to the Project Data Sphere platform is available to researchers affiliated with life science companies, hospitals and institutions, as well as independent researchers, at no cost and without requiring a research proposal.
 +
 +
* [https://www.smir.ch/ SICAS Medical Image Repository]([https://perma.cc/DMH2-EZQV perma.cc]). A place to store medical research data.
 +
 +
* [https://vivli.org/ Vivli]([https://perma.cc/2TSK-9EFB perma.cc]). From the Center for Global Clinical Research Data. The Vivli platform includes an independent data repository, in-depth search engine and a secure research environment.
  
 
== Multidisciplinary repositories ==
 
== Multidisciplinary repositories ==
Line 250: Line 294:
  
 
* [http://www.dans.knaw.nl/en/ Data Archiving and Networked Services]. Dutch research data in the humanities and social sciences.  From the [http://www.knaw.nl/ Royal Netherlands Academy of Arts and Sciences] (KNAW) and the [http://www.nwo.nl/ Netherlands Organisation for Scientific Research] (NWO).
 
* [http://www.dans.knaw.nl/en/ Data Archiving and Networked Services]. Dutch research data in the humanities and social sciences.  From the [http://www.knaw.nl/ Royal Netherlands Academy of Arts and Sciences] (KNAW) and the [http://www.nwo.nl/ Netherlands Organisation for Scientific Research] (NWO).
 +
 +
* [http://www.datacite.org.s3-website-eu-west-1.amazonaws.com/index.html DataCite] ([https://perma.cc/U3GN-AYBU perma.cc]).  DataCite is a leading global non-profit organisation that provides persistent identifiers (DOIs) for research data and other research outputs.
 +
 +
* [http://dataconservancy.org/community/ Data Conservancy]([https://perma.cc/J99T-HLC7 perma.cc]). Data Conservancy is devoted to developing institutional solutions for the challenges of data collection, preservation and re-use.
 +
 +
* [https://datahub.io/ DataHub]([https://perma.cc/6HF6-6RTL perma.cc]). There are thousands of datasets from financial market data and population growth to cryptocurrency prices.
  
 
* [http://dataone.org/ DataONE] DataONE is an international federation of data repositories containing earth observations data, including data from fields such as ecology, biology, evolution, and environmental sciences such as hydrology, oceanography, and atmospheric science.  DataONE is a federation with participation from hundreds of field stations, universities, and government agencies through the DataONE Member Nodes.
 
* [http://dataone.org/ DataONE] DataONE is an international federation of data repositories containing earth observations data, including data from fields such as ecology, biology, evolution, and environmental sciences such as hydrology, oceanography, and atmospheric science.  DataONE is a federation with participation from hundreds of field stations, universities, and government agencies through the DataONE Member Nodes.
Line 256: Line 306:
  
 
* [http://www.datadryad.org/ Dryad] Dryad is an international repository of data underlying scientific and medical publications, particularly data for which no specialized repository exists. All material in Dryad is associated with a scholarly publication. Most data in the repository are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted.  Dryad is a non-profit organization.  
 
* [http://www.datadryad.org/ Dryad] Dryad is an international repository of data underlying scientific and medical publications, particularly data for which no specialized repository exists. All material in Dryad is associated with a scholarly publication. Most data in the repository are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted.  Dryad is a non-profit organization.  
 +
 +
* [https://easy.dans.knaw.nl/ui/home EASY]([https://perma.cc/VK5M-VV64 perma.cc]). EASY offers sustainable archiving of research data and access to thousands of datasets.
  
 
* [http://datashare.is.ed.ac.uk/ Edinburgh DataShare] hosted by [http://www.ed.ac.uk/is/data-library Edinburgh University Data Library].  A repository for data produced by research at the [http://www.ed.ac.uk/ University of Edinburgh].  
 
* [http://datashare.is.ed.ac.uk/ Edinburgh DataShare] hosted by [http://www.ed.ac.uk/is/data-library Edinburgh University Data Library].  A repository for data produced by research at the [http://www.ed.ac.uk/ University of Edinburgh].  
  
 
* [http://www.ncbi.nlm.nih.gov/Database/ Entrez databases].  A directory of chemical, biochemical, biomedical, and medical databases from the U.S. [http://www.ncbi.nlm.nih.gov/ National Center for Biotechnology Information] of the [http://www.nih.gov/ National Institutes of Health].
 
* [http://www.ncbi.nlm.nih.gov/Database/ Entrez databases].  A directory of chemical, biochemical, biomedical, and medical databases from the U.S. [http://www.ncbi.nlm.nih.gov/ National Center for Biotechnology Information] of the [http://www.nih.gov/ National Institutes of Health].
 +
 +
* [https://eudat.eu/ EUDAT]([https://perma.cc/A4VG-3PCM perma.cc]). EUDAT offers heterogeneous research data management services and storage resources, supporting multiple research communities as well as individuals, through a geographically distributed, resilient network distributed across 15 European nations and data is stored alongside some of Europe’s most powerful supercomputers.
  
 
* [http://figshare.com/ FigShare]. Scientific publishing as it stands is an inefficient way to do science on a global scale. A lot of time and money is being wasted by groups around the world duplicating research that has already been carried out. FigShare allows you to share all of your data, negative results and unpublished figures. In doing this, other researchers will not duplicate the work, but instead may publish with your previously wasted figures, or offer collaboration opportunities and feedback on preprint figures.  
 
* [http://figshare.com/ FigShare]. Scientific publishing as it stands is an inefficient way to do science on a global scale. A lot of time and money is being wasted by groups around the world duplicating research that has already been carried out. FigShare allows you to share all of your data, negative results and unpublished figures. In doing this, other researchers will not duplicate the work, but instead may publish with your previously wasted figures, or offer collaboration opportunities and feedback on preprint figures.  
Line 268: Line 322:
  
 
* [http://kpbc.umk.pl/dlibra KPBC]. Regional academic repository for data in all fields. Poland
 
* [http://kpbc.umk.pl/dlibra KPBC]. Regional academic repository for data in all fields. Poland
 +
 +
* [https://msropendata.com/ Microsoft Research Open Data]([https://perma.cc/F59P-ANLM perma.cc]). A collection of free datasets from Microsoft Research to advance state-of-the-art research in areas such as natural language processing, computer vision, and domain specific sciences. Download or copy directly to a cloud-based Data Science Virtual Machine for a seamless development experience.
  
 
* [http://www.occ-data.org/ Open Commons Consortium (OCC)].  The OCC is a not for profit that manages and operates cloud computing and data commons infrastructure to support scientific, medical, health care and environmental research. OCC members span the globe and include over 30 universities, companies, government agencies and national laboratories.   
 
* [http://www.occ-data.org/ Open Commons Consortium (OCC)].  The OCC is a not for profit that manages and operates cloud computing and data commons infrastructure to support scientific, medical, health care and environmental research. OCC members span the globe and include over 30 universities, companies, government agencies and national laboratories.   
Line 282: Line 338:
  
 
* [http://www.science3point0.com/opendata/index.php Science 3.0 Open Data].  A repository for RDF datasets in the public domain, in any field.  From [http://www.science3point0.com/ Science 3.0].
 
* [http://www.science3point0.com/opendata/index.php Science 3.0 Open Data].  A repository for RDF datasets in the public domain, in any field.  From [http://www.science3point0.com/ Science 3.0].
 +
 +
* [https://figshare.com/articles/Scientific_Data_recommended_repositories_June_2015/1434640 Scientific Data recommended repositories]([https://perma.cc/FLG9-BB7Z perma.cc]). Spreadsheet listing data repositories that are recommended by Scientific Data (Springer Nature) as being suitable for hosting data associated with peer-reviewed articles. Please see the repository list on Scientific Data's website for the most up to date list.
 +
 +
* [http://site.uit.no/trolling/about/ Tromsø Repository of Language and Linguistics (TROLLing)]([https://perma.cc/9BCZ-J42Y perma.cc]). TROLLing is designed as an archive of linguistic data and statistical code. The archive is open access, which means that all information is available to to everyone. All postings are accompanied by searchable metadata that identify the researchers, the languages and linguistic phenomena involved, the statistical methods applied, and scholarly publications based on the data (where relevant).
  
 
* [http://repository.usu.ac.id/ USU Repository] University of Sumatera Utara, Medan, Indonesia.
 
* [http://repository.usu.ac.id/ USU Repository] University of Sumatera Utara, Medan, Indonesia.
  
 
* [http://repository.up.ac.za/upspace/ UPSpace] University of Pretoria Research Repository, South Africa.
 
* [http://repository.up.ac.za/upspace/ UPSpace] University of Pretoria Research Repository, South Africa.
 +
 +
* [https://webscope.sandbox.yahoo.com/?guccounter=2 Webscope]([https://perma.cc/AND7-EZGJ perma.cc]). The Yahoo Webscope Program is a reference library of interesting and scientifically useful datasets for non-commercial use by academics and other scientists. All datasets have been reviewed to conform to Yahoo's data protection standards, including strict controls on privacy.
 +
 +
* [https://www.zenodo.org/ Zenodo]([https://perma.cc/U4VW-TXBN perma.cc]). All research outputs from across all fields of research.
  
 
== Physics ==
 
== Physics ==
Line 292: Line 356:
  
 
* [http://bodr.sf.net/ Blue Obelisk Data Repository]. Repository of isotope masses, under MIT license. From the [http://www.blueobelisk.org/ Blue Obelisk]. Described in [http://dx.doi.org/10.1021/ci050400b 10.1021/ci050400b].
 
* [http://bodr.sf.net/ Blue Obelisk Data Repository]. Repository of isotope masses, under MIT license. From the [http://www.blueobelisk.org/ Blue Obelisk]. Described in [http://dx.doi.org/10.1021/ci050400b 10.1021/ci050400b].
 +
 +
* [https://library.web.cern.ch/library/rpp/ CERN Scientific Information] Online particle physics data and information
  
 
* [http://hepdata.cedar.ac.uk/submittingdata HEP Data] The data comprise total and differential cross sections, structure functions, fragmentation functions, distribuitions of jet measures, polarisations, etc... from a wide range of interactions.
 
* [http://hepdata.cedar.ac.uk/submittingdata HEP Data] The data comprise total and differential cross sections, structure functions, fragmentation functions, distribuitions of jet measures, polarisations, etc... from a wide range of interactions.
  
* [https://library.web.cern.ch/library/rpp/ CERN Scientific Information] Online particle physics data and information
+
* [https://lambda.gsfc.nasa.gov/ LAMBDA]([https://perma.cc/44NM-FH9E perma.cc]). LAMBDA is a part of NASA's [http://heasarc.gsfc.nasa.gov/ High Energy Astrophysics Science Archive Research Center] (HEASARC).
  
 
* [http://www.nist.gov/pml/data/asd.cfm Nist Atomic Spectra Database] The Atomic Spectra Database (ASD) contains data for radiative transitions and energy levels in atoms and atomic ions.  Data are included for observed transitions of 99 elements and energy levels of 56 elements.
 
* [http://www.nist.gov/pml/data/asd.cfm Nist Atomic Spectra Database] The Atomic Spectra Database (ASD) contains data for radiative transitions and energy levels in atoms and atomic ions.  Data are included for observed transitions of 99 elements and energy levels of 56 elements.
Line 314: Line 380:
  
 
* [http://www.esds.ac.uk/ Economic and Social Science Data Service]. From the [http://www.data-archive.ac.uk/ UK Data Archive (UKDA)] and [http://www.iser.essex.ac.uk/ Institute for Social and Economic Research (ISER)], University of Essex; [http://www.mimas.ac.uk/ Manchester Information and Associated Services (MIMAS)], and the [http://www.ccsr.ac.uk/ Cathie Marsh Centre for Census and Survey Research (CCSR)], University of Manchester. Access to data requires registration.
 
* [http://www.esds.ac.uk/ Economic and Social Science Data Service]. From the [http://www.data-archive.ac.uk/ UK Data Archive (UKDA)] and [http://www.iser.essex.ac.uk/ Institute for Social and Economic Research (ISER)], University of Essex; [http://www.mimas.ac.uk/ Manchester Information and Associated Services (MIMAS)], and the [http://www.ccsr.ac.uk/ Cathie Marsh Centre for Census and Survey Research (CCSR)], University of Manchester. Access to data requires registration.
 +
 +
* [https://www.ebi.ac.uk/ena European Nucleotide Archive]([https://perma.cc/ZE72-PECM perma.cc]). The European Nucleotide Archive (ENA) provides a comprehensive record of the world's nucleotide sequencing information, covering raw sequencing data, sequence assembly information and functional annotation.
  
 
* [http://www.icpsr.umich.edu/icpsrweb/ICPSR/ ICPSR] (Inter-University Consortium for Political and Social Research).  At the University of Michigan.
 
* [http://www.icpsr.umich.edu/icpsrweb/ICPSR/ ICPSR] (Inter-University Consortium for Political and Social Research).  At the University of Michigan.
  
 
* [http://www.icpsr.umich.edu/icpsrweb/NACJD/index.jsp National Archive of Criminal Justice Data] holds over 700 data collections relating to criminal justice.
 
* [http://www.icpsr.umich.edu/icpsrweb/NACJD/index.jsp National Archive of Criminal Justice Data] holds over 700 data collections relating to criminal justice.
 +
 +
* [http://nomad-repository.eu/ NOMAD Repository]([https://perma.cc/SU9A-YTRQ perma.cc]). Host, organize, and share materials data.
 +
 +
* [https://www.openicpsr.org/openicpsr/ openICPRS]([https://perma.cc/2QEZ-84DS perma.cc]). openICPSR is a great place to share and store your social and behavioral science research data. Your data will be preserved as-is and be available to data users at no cost.
 +
 +
* [https://qdr.syr.edu/ Qualitative Data Repository]([https://perma.cc/9GGP-D7PB perma.cc]). QDR curates, stores, preserves, publishes, and enables the download of digital data generated through qualitative and multi-method research in the social sciences.
  
 
* [http://ropercenter.cornell.edu/ Roper Center for Public Opinion Research] data from surveys of public opinion from the 1930s to the present.
 
* [http://ropercenter.cornell.edu/ Roper Center for Public Opinion Research] data from surveys of public opinion from the 1930s to the present.

Revision as of 18:50, 12 February 2020

Oad2.jpeg This list is part of the Open Access Directory.

  • This is a list of repositories and databases for open data.
  • Please annotate the entries to indicate the hosting organization, scope, licensing, and usage restrictions (if any). If a repository is open in some respects but not others, please include it with an annotation rather than exclude it.
  • If you're not sure whether a given dataset or data collection is open, post your query to Is It Open Data?
  • Related lists in OAD: Disciplinary repositories (primarily for texts, not data).
  • For news about data repositories, including some newly launched repositories not yet listed here, follow the oa.repositories.data tag of the Open Access Tracking Project.
  • See also:

Archaeology

  • Also see Social sciences.
  • Fasti Online . Subdivided in Excavation, Restauration and Survey.

Astronomy

  • Also see Physics.

Biology

  • Also see BCO-DMO, Marine Biology data, listed with Marine Sciences repositories.
  • Also see DataONE, Entrez databases, KNB, and PANGAEA, listed under Multidisciplinary repositories.
  • The Cell: An Image Library Images of all cell types from all organisms, including intracellular structures and movies or animations demonstrating functions. This project relies upon the cell biology community to populate the library. The Cell: An Image Library™ is a freely accessible, easy-to-search, public repository of reviewed and annotated images, videos, and animations of cells from a variety of organisms, showcasing cell architecture, intracellular functionalities, and both normal and abnormal processes. The purpose of this database is to advance research, education, and training, with the ultimate goal of improving human health.
  • Database of Virulence Factors in Fungal Pathogenes (DFVF)(perma.cc). The database is expected to greatly stimulate and facilitate further studies in fungal pathogens; both experimental biologists and computational biologists can use the database and/or the predicted virulence factors to guide their search for new virulence factors and/or discovery of new pathogen-host interaction mechanisms in fungi.
  • dpSNP(perma.cc). dbSNP contains human single nucleotide variations, microsatellites, and small-scale insertions and deletions along with publication, population frequency, molecular consequence, and genomic and RefSeq mapping information for both common variations and clinical mutations.
  • dbVar(perma.cc). dbVar is NCBI's database of human genomic Structural Variation — large variants >50 bp including insertions, deletions, duplications, inversions, mobile elements, translocations, and complex variants.
  • Dryad Dryad is an international repository of data underlying scientific and medical publications, particularly data for which no specialized repository exists. All material in Dryad is associated with a scholarly publication. Most data in the repository are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted. Dryad is a non-profit organization.
  • Gene Expression Omnibus High-throughput functional genomic data, including all array-based applications and some high-throughput sequencing data.
  • Human Protein Atlas(perma.cc). All the data in the knowledge resource is open access to allow scientists both in academia and industry to freely access the data for exploration of the human proteome.
  • MGnify(perma.cc). MGnify offers an automated pipeline for the analysis and archiving of microbiome data to help determine the taxonomic diversity and functional & metabolic potential of environmental samples.
  • National Biological Information Infrastructure A broad, collaborative program to provide increased access to data and information on the nation's biological resources. The NBII links diverse, high-quality biological databases, information products, and analytical tools maintained by NBII partners and other contributors in government agencies, academic institutions, non-government organizations, and private industry. (Note: In the President's budget for Fiscal Year 2012 the repository was terminated.)
  • Planet A network of European Plant Databases
  • The Universal Protein Resource (UniProt) is a comprehensive resource for protein sequence and annotation data. The UniProt databases are the UniProt Knowledgebase (UniProtKB), the UniProt Reference Clusters (UniRef), and the UniProt Archive (UniParc). The UniProt Metagenomic and Environmental Sequences (UniMES) database is a repository specifically developed for metagenomic and environmental data.
  • uBio(perma.cc). uBio uses names and taxonomic intelligence to manage information about organisms.

Chemistry

  • Also see BCO-DMO, Marine Biology data, listed with Marine Sciences repositories.
  • Also see Entrez databases, listed under Multidisciplinary repositories.
  • Cambridge Structural Database The CCDC is a non-profit, charitable Institution whose objectives are the general advancement and promotion of the science of chemistry and crystallography for the public benefit.
  • ChemSpider. Hosted by the Royal Society of Chemistry.
  • ChemSynthesis. A database of chemicals and their physical properties.
  • eCrystals. From the Southampton Chemical Crystallography Group and the EPSRC UK National Crystallography Service.

Computer Science

  • CiteSeerX provides its databases of nearly 2 million documents and the associated texts and pdfs for research.
  • FreeStatistics of Irreproducible Research(perma.cc). The purpose of this project is to facilitate the creation, maintenance, and permanent storage of statistical computation objects that empower authors to publish reproducible and reusable research (in the form of a Compendium) through a series of web services.
  • GitHub keeps your public and private code available, secure, and backed up.
  • Google Code Project Hosting Project Hosting on Google Code provides a free collaborative development environment for open source projects. Each project comes with its own member controls, Subversion/Mercurial repository, issue tracker, wiki pages, and downloads section. Our project hosting service is simple, fast, reliable, and scalable, so that you can focus on your own open source development.
  • Launchpad can host your project’s source code using the Bazaar version control system. We also import over 2000 CVS, SVN, Git and Mercurial projects, so you can use Bazaar with those too.
  • ReproZip!(perma.cc). ReproZip can automatically pack your research along with all necessary data files, libraries, environment variables and options into a self-contained bundle. Then ReproZip can use that bundle to automatically set up the same original environment so anybody can reproduce the research on a different machine, without tracking down and installing the dependencies, or even having to run the same operating system.
  • SourceForge 2.7 million developers create powerful software in over 260,000 projects. Our popular directory connects more than 46 million consumers with these open source projects and serves more than 2,000,000 downloads a day. SourceForge is where open source happens.
  • SNAP Stanford Large Network Dataset Collection. The SNAP library is being actively developed since 2004 and is organically growing as a result of our research pursuits in analysis of large social and information networks. Largest network we analyzed so far using the library was the Microsoft Instant Messenger network from 2006 with 240 million nodes and 1.3 billion edges.
  • KONECT (the Koblenz Network Collection) is a project to collect large network datasets of all types in order to perform research in network science and related fields, collected by the Institute of Web Science and Technologies at the University of Koblenz–Landau.

Energy

Environmental sciences

  • Also see BCO-DMO, Marine Biology data, listed with Marine Sciences repositories.
  • Also see DataONE, KNB, and PANGAEA, listed under Multidisciplinary repositories.
  • Also see Dryad, listed with Biology repositories.
  • The Marine Geoscience Data System (MGDS) The Marine Geoscience Data System (MGDS) provides access to data portals for the NSF-supported Ridge 2000 and MARGINS programs, the Antarctic and Southern Ocean Data Synthesis, the Global Multi-Resolution Topography Synthesis, and Seismic Reflection Field Data Portal.
  • Polar Data Catalogue A primarily Canadian archive of free RADARSAT imagery as well as Arctic, Antarctic, and other cryospheric data sets covering a range of disciplines, from natural sciences and policy to health and social sciences.
  • Socioeconomic Data and Applications Center (SEDAC) specializes in spatial data and services in support of human-environment research and applications, in the context of NASA’s Earth science mission and the overall U.S. Global Change Research Program.

Geology

  • Also see PANGAEA, listed under Multidisciplinary repositories.
  • IRIS (Incorporated Research Institutions for Seismology). From 100+ US universities and the National Science Foundation.

Geosciences and geospatial data

  • Also see DataONE and PANGAEA, listed under Multidisciplinary repositories.
  • EarthChem Library(perma.cc). The EarthChem Library is a data repository that archives, publishes and makes accessible data and other digital content from geoscience research (analytical data, data syntheses, models, technical reports, etc).
  • GeoNames. A database of placenames, under a CC-BY license. Founded by Marc Wick.
  • The Geosciences Network (GEON) project is a collaboration among a dozen PI institutions and a number of other partner projects, institutions, and agencies to develop cyberinfrastructure in support of an environment for integrative geoscience research. GEON is funded by the NSF Information Technology Research (ITR) program.
  • National Space Science Data Center serves as the permanent archive for NASA space science mission data. "Space science" means astronomy and astrophysics, solar and space plasma physics, and planetary and lunar science. As permanent archive, NSSDC teams with NASA's discipline-specific space science "active archives" which provide access to data to researchers and, in some cases, to the general public.
  • OpenTopography(perma.cc). OpenTopography facilitates community access to high-resolution, Earth science-oriented, topography data, and related tools and resources.
  • Polar Data Catalogue A primarily Canadian archive of free RADARSAT imagery as well as Arctic, Antarctic, and other cryospheric data sets covering a range of disciplines, from natural sciences and policy to health and social sciences.
  • ShareGeo. Integrating the older GRADE (Geospatial Repository for Academic Deposit and Extraction) repository. From EDINA.

Linguistics

  • See the 40+ members of the Open Language Archives Community (OLAC).
  • TROLLing. Hosted by UiT. TROLLing "is designed as an archive of linguistic data and statistical code. The archive is open access, which means that all information is available to to everyone. All postings are accompanied by searchable metadata that identify the researchers, the languages and linguistic phenomena involved, the statistical methods applied, and scholarly publications based on the data (where relevant). Linguists worldwide are invited to post datasets and statistical models used in linguistic research."

Marine sciences

  • Also see DataONE and PANGAEA, listed under Multidisciplinary repositories.
  • BCO-DMO. The Biological and Chemical Oceanography Data Management Office, provides access to data sets contributed by investigators funded by the Biological and Chemical Oceanography sections of the US National Science Foundation (NSF).
  • SEAONE - Sea Open Scientific Data Publication(perma.cc). SEANOE (SEA scieNtific Open data Edition) is a publisher of scientific data in the field of marine sciences. Data published by SEANOE are available free. They can be used in accordance with the terms of the Creative Commons license selected by the author of data.

Medicine

  • Also see Entrez databases, listed under Multidisciplinary repositories.
  • Dryad Dryad is an international repository of data underlying scientific and medical publications, particularly data for which no specialized repository exists. All material in Dryad is associated with a scholarly publication. Most data in the repository are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted. Dryad is a non-profit organization.
  • FlowRepository(perma.cc). FlowRepository is a database of flow cytometry experiments where you can query and download data collected and annotated according to the MIFlowCyt standard.
  • The Health and Medical Care Archive (HMCA) is the data archive of the Robert Wood Johnson Foundation (RWJF), the largest philanthropy devoted exclusively to health and health care in the United States. Operated by the Inter-university Consortium for Political and Social Research (ICPSR) at the University of Michigan, HMCA preserves and disseminates data collected by selected research projects funded by the Foundation and facilitates secondary analyses of the data. The data collections in HMCA include surveys of health care professionals and organizations, investigations of access to medical care, surveys on substance abuse, and evaluations of innovative programs for the delivery of health care. Our goal is to increase understanding of health and health care in the United States through secondary analysis of RWJF-supported data collections.
  • MIRAGE (Middlesex medical Image Repository with a CBIR ArchivinG Environment). From JISC and Middlesex University.
  • Project Data Sphere, LLC, is a repository to broadly share, integrate and analyze historical, de-identified, patient-level data from academic and industry cancer Phase II-III clinical trials. Access to the Project Data Sphere platform is available to researchers affiliated with life science companies, hospitals and institutions, as well as independent researchers, at no cost and without requiring a research proposal.
  • Vivli(perma.cc). From the Center for Global Clinical Research Data. The Vivli platform includes an independent data repository, in-depth search engine and a secure research environment.

Multidisciplinary repositories

  • Also see Social Sciences.
  • Also see BCO-DMO, Marine Biology data, listed with Marine Sciences repositories.
  • DataCite (perma.cc). DataCite is a leading global non-profit organisation that provides persistent identifiers (DOIs) for research data and other research outputs.
  • Data Conservancy(perma.cc). Data Conservancy is devoted to developing institutional solutions for the challenges of data collection, preservation and re-use.
  • DataHub(perma.cc). There are thousands of datasets from financial market data and population growth to cryptocurrency prices.
  • DataONE DataONE is an international federation of data repositories containing earth observations data, including data from fields such as ecology, biology, evolution, and environmental sciences such as hydrology, oceanography, and atmospheric science. DataONE is a federation with participation from hundreds of field stations, universities, and government agencies through the DataONE Member Nodes.
  • Dryad Dryad is an international repository of data underlying scientific and medical publications, particularly data for which no specialized repository exists. All material in Dryad is associated with a scholarly publication. Most data in the repository are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted. Dryad is a non-profit organization.
  • EASY(perma.cc). EASY offers sustainable archiving of research data and access to thousands of datasets.
  • EUDAT(perma.cc). EUDAT offers heterogeneous research data management services and storage resources, supporting multiple research communities as well as individuals, through a geographically distributed, resilient network distributed across 15 European nations and data is stored alongside some of Europe’s most powerful supercomputers.
  • FigShare. Scientific publishing as it stands is an inefficient way to do science on a global scale. A lot of time and money is being wasted by groups around the world duplicating research that has already been carried out. FigShare allows you to share all of your data, negative results and unpublished figures. In doing this, other researchers will not duplicate the work, but instead may publish with your previously wasted figures, or offer collaboration opportunities and feedback on preprint figures.
  • KPBC. Regional academic repository for data in all fields. Poland
  • Microsoft Research Open Data(perma.cc). A collection of free datasets from Microsoft Research to advance state-of-the-art research in areas such as natural language processing, computer vision, and domain specific sciences. Download or copy directly to a cloud-based Data Science Virtual Machine for a seamless development experience.
  • Open Commons Consortium (OCC). The OCC is a not for profit that manages and operates cloud computing and data commons infrastructure to support scientific, medical, health care and environmental research. OCC members span the globe and include over 30 universities, companies, government agencies and national laboratories.
  • Open Science Data Cloud (OSDC). The OSDC is a data science ecosystem in which researchers can house and share their own scientific data, access complementary public datasets, build and share customized virtual machines with whatever tools necessary to analyze their data, and perform the analysis to answer their research questions. It is a one-stop shop for making scientific research faster and easier.
  • Open Science Framework (OSF) Open Science Framework serves as a scholarly commons for documentation, files, collaboration, and connecting to services for research outputs.
  • Scientific Data recommended repositories(perma.cc). Spreadsheet listing data repositories that are recommended by Scientific Data (Springer Nature) as being suitable for hosting data associated with peer-reviewed articles. Please see the repository list on Scientific Data's website for the most up to date list.
  • Tromsø Repository of Language and Linguistics (TROLLing)(perma.cc). TROLLing is designed as an archive of linguistic data and statistical code. The archive is open access, which means that all information is available to to everyone. All postings are accompanied by searchable metadata that identify the researchers, the languages and linguistic phenomena involved, the statistical methods applied, and scholarly publications based on the data (where relevant).
  • UPSpace University of Pretoria Research Repository, South Africa.
  • Webscope(perma.cc). The Yahoo Webscope Program is a reference library of interesting and scientifically useful datasets for non-commercial use by academics and other scientists. All datasets have been reviewed to conform to Yahoo's data protection standards, including strict controls on privacy.
  • Zenodo(perma.cc). All research outputs from across all fields of research.

Physics

  • Also see Astronomy.
  • HEP Data The data comprise total and differential cross sections, structure functions, fragmentation functions, distribuitions of jet measures, polarisations, etc... from a wide range of interactions.
  • Nist Atomic Spectra Database The Atomic Spectra Database (ASD) contains data for radiative transitions and energy levels in atoms and atomic ions. Data are included for observed transitions of 99 elements and energy levels of 56 elements.

Social sciences

  • Also see Multidisciplinary repositories.
  • Databrary A repository for sharing and reusing research video data and related metadata in the developmental and learning sciences. Hosted at New York University with support from The Pennsylvania State University.
  • European Nucleotide Archive(perma.cc). The European Nucleotide Archive (ENA) provides a comprehensive record of the world's nucleotide sequencing information, covering raw sequencing data, sequence assembly information and functional annotation.
  • ICPSR (Inter-University Consortium for Political and Social Research). At the University of Michigan.
  • openICPRS(perma.cc). openICPSR is a great place to share and store your social and behavioral science research data. Your data will be preserved as-is and be available to data users at no cost.
  • Qualitative Data Repository(perma.cc). QDR curates, stores, preserves, publishes, and enables the download of digital data generated through qualitative and multi-method research in the social sciences.