Laying a foundation for meaningful data sharing

An interview with Dr. Pradip K. Bera

For Dr. Pradip K. Bera organization and planning are both the major challenge and the key to effective data sharing.

Bera is a postdoctoral research associate in the Department of Mechanical Engineering at the University of Wisconsin-Madison, where he works at the intersection of soft condensed matter, nonequilibrium physics, and biological physics. He’s particularly interested in how collective material properties emerge in systems ranging from polymers and colloids to living tissues. 

In July, Bera and colleagues published the dataset “DECMA-1 main data set for: Shape-independent fluidity in epithelial cell monolayers” with Dryad, and posted two related scripts available on Zenodo. The work challenges a long-held understanding of how fluidity impacts cell shape, with implications for everything from wound healing, to development, to disease.

“The prevailing view has been that tissue fluidity and cell geometry are closely interlinked through the competition between cell-cell adhesion and cortical tension, such that changes in fluidity are accompanied by changes in cell shape,” he explained. “However, we showed that epithelial cell monolayers can become substantially more fluid when cell-cell adhesion is reduced, even though the average cell shape remains nearly unchanged. Our results suggest that cell-cell adhesion has two independent contributions to tissue fluidity: one through cell geometry and the other through intercellular dissipation.”

Value demands access

The major benefits of data sharing are reproducibility, transparency, reuse, and the possibility of generating new science from existing experiments, Bera said.

The researchers chose to make their data public in part because of its importance to the field. 

“These datasets are particularly valuable because they contain multiple measurements performed in parallel experiments, requiring experimental capabilities that may not be readily available to many researchers,” Bera said. “Making the data publicly available allows other researchers to independently examine our results, apply alternative analysis approaches, and build upon the data in their future work. I also believe that sharing well-organized research data strengthens reproducibility and increases the long-term scientific value of an experiment.”

I believe that sharing well-organized research data strengthens reproducibility and increases the long-term scientific value of an experiment.”

He anticipates researchers may use the dataset to reproduce the analyses, compare new theoretical models with experimental measurements, or develop and benchmark alternative image-analysis and tissue-mechanics methods. “The dataset may also enable researchers to explore new questions that we did not originally consider when designing the experiments,” Bera said.

The team selected Dryad for their dataset “because it is a well-established repository specifically designed for research data, offering long-term preservation, a permanent DOI, and good integration with scholarly publishing. Its professional data curation and the ability to make our dataset easily discoverable, accessible, and citable were particularly important to us.”

Organization and planning provide the key

“The main effort was not uploading the files themselves, but organizing and documenting the data so that someone outside our group could understand what each file represented and how it related to the published figures and analyses,” Bera said. “This process required additional time, but it also helped us make the dataset more structured, understandable, and reusable.”

Undergoing the review process at Dryad “reinforced how valuable it is to organize and document data early in a research project, rather than waiting until the publication stage.”

That’s Bera’s main advice to fellow researchers preparing to share their data: “Organize the data with a new researcher in mind: use meaningful file names, preserve raw data where possible, clearly describe processing steps, and include enough metadata to connect the files to the publication. Starting this organization early in the research project, rather than waiting until the end, can save considerable time and make the data more useful to others.”

Another challenge: measuring impact and allocating credit. “One important challenge is tracking how publicly available datasets are downloaded and reused, and ensuring that the original datasets are properly cited when they contribute to new publications,” said Bera.

Data sharing is becoming normal. Hopefully, that’s just the beginning

Data sharing is increasingly becoming part of the normal publication workflow rather than an optional addition after a study is completed. Researchers are also becoming more aware that datasets themselves can be valuable and citable scientific outputs., Bera said.

That’s something Bera has discovered first hand. He has accessed publicly available datasets and analysis code to better understand published methodologies and compare published results with his own. Here, too, organization is key. “My experience has been positive, particularly when the data and code were well organized and accompanied by clear documentation.”

Looking further ahead, Bera hopes to see greater emphasis on the useful and reusable primary research outputs. 

“I hope publications increasingly become connected packages of papers, data, code, and documentation, making scientific results easier to verify, reproduce, and build upon. At the same time, rigorous peer review and careful scientific interpretation should remain central, because open data are most valuable when accompanied by clear context and responsible analysis.”

Share your data story. Have you shared open data or re-used others’ publicly-available data? Share your story with us at partnerships [at] datadryad.org.

Feedback and questions are always welcome, to hello@datadryad.org

To keep in touch with the latest updates from Dryad, follow us on LinkedIn, Mastodon, and Bluesky and subscribe to our quarterly newsletter.