Dryad and CEDAR: Supporting more research communities with generalist infrastructure

Dryad is a leading publisher of research data in all its varied forms and formats across disciplines. We work with thousands of authors each year to prepare data publications according to FAIR standards, and house over 70,000 highly-discoverable preserved research datasets. Our generalist approach to curating and publishing data allows us to efficiently and effectively support over 1,000 authors each month as well as a growing community of institutions and publishing organizations that rely on our service. 

But we also recognize the critical importance of upholding discipline-specific standards that reflect best practices in individual fields — both to promote reproducibility and research integrity, and to drive discovery and reuse. We encourage researchers to share appropriate data with specialist repositories where they are available.

Where specialist solutions do not exist, the Center for Expanded Data Annotation and Retrieval (CEDAR) at Stanford University has developed an embeddable editor that integrates with repositories like Dryad, allowing researchers to describe research data in a discipline-appropriate way, while leveraging extensive, established repository infrastructure. The CEDAR Embeddable Editor overlays broadly-accepted discipline-specific metadata templates on top of Dryad’s default FAIR metadata schema, maintained by DataCite. CEDAR templates are developed once, and used across repository platforms without modification for maximum interoperability.

As researchers move through the Dryad data submission process, we automatically scan title, abstract, and other metadata to assign keywords. Depending on the keywords identified, our system surfaces optional supplementary metadata templates that may be relevant to that particular dataset.

At publication CEDAR metadata is included as JSON (DisciplineSpecificMetadata.json) in the dataset files. It can be opened in a read-only version of the CEDAR embeddable editor on the dataset landing page.

Our work with the NIH’s Human BioMolecular Atlas Program (HuBMAP) provides just one example. HuBMAP offers more than 35 different templates for specific use cases for annotating data for biological assays, including RNAseq, FACS, histology, and ATACseq. Since implementing the integration, more than 65 Dryad datasets have implemented enhanced metadata.

The big picture

One of the larger impediments to Open Science is the fractured nature of open infrastructure. Thousands of repositories, small nonprofit organizations, software providers, journals, databases, reporting and taxonomy standards, and metadata schemas all contribute to open and interoperable research, but pull in too many different directions. This complexity causes inefficiency and creates cost, hampering the ability of individual services to establish sustainability as well as the more widespread embrace of Open Science practices.

Tools like CEDAR model one way to limit some of this complexity. These flexible but interoperable metadata overlays help individual research communities collect discipline-specific semantically rich metadata while avoiding building and maintaining additional specialist infrastructure. Discipline-specific metadata, as supported by Dryad and CEDAR, enable much more effective dataset search, facilitating reuse and potentially leading to new discoveries. Together, we can benefit from economies of scale and address the needs of individual research communities; we can be both scalable and specific. 

Add a discipline-specific metadata template to Dryad

Does your researcher community have a broadly accepted metadata schema? Enriched metadata holds the potential to advance interdisciplinary research and accelerate scientific discoveries. Get in touch to start a conversation about implementing a metadata template for your community.  

This feature is powered by the CEDAR Embeddable Editor. Initial work for this project was funded by the U.S. National Science Foundation, award 2134956.

Feedback and questions are always welcome, to hello@datadryad.org

To keep in touch with the latest updates from Dryad, follow us on LinkedIn, Mastodon, and Bluesky and subscribe to our quarterly newsletter.