Category Archives: Data Curation

Joel Meyer on sharing environmental health data for the public good

500 Datasets, 5 Researchers, 5 StoriesThis story is part of a fall 2026 series celebrating 500 published datasets in the Duke Research Data Repository (RDR) by spotlighting 5 Duke researchers who make their work more transparent, discoverable, and reusable.

Joel Meyer, Duke ProfessorJoel Meyer is a Sally Kleberg Distinguished Professor in the Nicholas School of the Environment, Director of the Integrated Toxicology and Environmental Health Program, and the Associate Director of Undergraduate Studies. Dr. Meyer is also a Principal Investigator within the Superfund Research Center. The Superfund Research Center focuses on early, low-dose exposures to environmental contaminants and their developmental impacts, changes usually only evident later in life (NIEHS grant, P42ES010356). The Superfund Research Center has a collection within the Duke Research Data Repository to highlight the public data generated by the Center’s investigators and demonstrate compliance with NIH data management and sharing policy. The Collection currently contains almost 30 datasets; the Center also has shared data in disciplinary repositories linked on their website.

Dr. Meyer shared with us his thoughts on sharing environmental health data within the Duke Research Data Repository (RDR).

Meyer Lab
The Meyer Lab, April 2026

Tell us a little about the research and the kinds of data shared through the RDR. What should someone outside your field know about them?

My group studies the effects of pollutants on health, largely using cells in culture and the small nematode (worm) Caenorhabditis elegans. We gather a wide range of data, from molecular impacts (e.g., does the chemical cause DNA damage or oxidative stress) to cellular effects (e.g., does the chemical cause neurons to degenerate or malfunction) to whole-organism effects (e.g., does the chemical inhibit growth or reproduction). We often generate datasets with many thousands of data points per experiment.

Why do you choose to share your research data?

I think it is important that other researchers, as well as other interested people in general, be able to access the data for themselves. They might want to analyze it in a different way, or see if they agree with our analysis, or combine the data with other data for larger-scale studies. I also think it is our responsibility to make this data generally available since most of it was generated with taxpayer money. Finally, I hope that this also promotes a sense of scientific transparency.

What do you hope other researchers, or other audiences, might be able to do with the data you’ve shared?

Superfund Research Center Team
Superfund Research Center Team

I hope that it will be useful to regulators who are trying to protect the public from negative health effects of pollutants. I also hope that the data may be of use to community members, nonprofit organizations, and other researchers. In fact, we recently published a paper on how to do academic research in a way that is broadly useful (open access!), and the use of data repositories, like the RDR, is one of the things we highlight.

What has your experience been like working with the Research Data Repository team? Is there anything about the publishing or curation process that has been particularly useful to you?

It has been great! The team has been very fast to respond to questions and provide feedback, and very helpful regarding how to describe and structure the data for maximum utility. Frankly, they and the process have improved our own processes for organizing and storing data and metadata.

What would you say to another Duke researcher who is considering sharing their data?

Do it!

Explore Dr Meyer’s and other Duke researcher’s datasets published in the Superfund Research Center RDR Collection.

If your research grant or project would like to create a Collection within the RDR to more easily highlight datasets, feel free to contact us!

The RDR team thanks Dr. Meyer for his participation and supporting data publishing at Duke. For questions about the RDR contact: datamanagement@duke.edu

 

 

 

Patrick Charbonneau on why he shares his data with the world

This story is part of a fall 2026 series celebrating 500 published datasets in the Duke Research Data Repository by spotlighting 5 Duke researchers who make their work more transparent, discoverable, and reusable.

Dr. Patrick Charbonneau maintains appointments in both Chemistry and Physics. Dr. Charbonneau published the 500th dataset within the Duke Research Data Repository (RDR) and is the top depositor. He shared with us his thoughts on data sharing and the use of the RDR.

Dr. Patrick CharbonneauTell us a little about the research and the kinds of data shared through the RDR. What should someone outside your field know about them?

My group works at the intersection of theoretical chemistry and physics, using statistical mechanics and computational methods to understand complex systems. The datasets we share typically contain the numerical data behind published figures, along with code and documentation needed to understand or reproduce the results.

Why do you choose to share your research data? Has your thinking about data sharing changed over time?

Early in my career, archiving meant asking students and postdocs, usually as they were leaving, to organize their files on a departmental drive. I eventually realized that this approach was both stressful and fragile. We therefore shifted to preparing data, code, and documentation alongside each manuscript, so that sharing becomes part of publishing rather than something done retrospectively. An unexpected benefit is that preparing a dataset for others to use provides an additional check on the research itself.

What do you hope other researchers, or other audiences, might be able to do with the data you’ve shared?

Dr. Charbonneau with then graduate student, Lin Fu, in 2012

I want our publications to remain reusable without having to track down a former student, postdoc, or me(!) years later. Surprisingly, this happened twice just this summer for papers from before we routinely deposited datasets; I have since deposited those data. But I also like that we cannot entirely predict what the data might eventually be useful for. A dataset created to answer one question may later become useful for comparison, methodological development, teaching, or a question we had never considered.

What has your experience been like working with the Duke Research Data Repository team? Is there anything about the publishing or curation process that has been particularly useful to you?

The RDR team has made data publication remarkably easy to incorporate into our research workflow. The curatorial oversight is particularly valuable: having another set of eyes on the organization, documentation, and long-term usability of a dataset helps turn files associated with a paper into a durable research product. Over time, working with the RDR has made depositing data simply part of what we consider completing a publication.

What would you say to another Duke researcher who is considering sharing their data?

Start small. There is no need to solve every possible question about what should be preserved before making a first deposit. Sharing even the numerical values behind published figures is a useful beginning, and once the process becomes part of the publication workflow, expanding what you share becomes much easier. In my experience, the cost is small, especially with the support of the RDR team, while the benefits for transparency, preservation, and future reuse can be substantial. Also, LLMs now make writing useful README files easier than ever.

A tray of pieces of sucre à la crème
A tray of pieces of sucre à la crème, the data from the associated article is published in the RDR

Has anything interesting or unexpected happened as a result of sharing data?

In recent years, I started publishing articles on the history of science and related questions. Surprisingly, the datasets I deposited to accompany those articles—oral history interviews, textual corpora, and archival documents—are among the most consulted of all my deposits. I certainly didn’t expect my humanities datasets to be the popular ones!

Charbonneau’s humanities datasets: Data from: Sucre à la crème: Origin and Trajectory of an Authentic Québec Confection; Data from: Pralines des Voyageurs: An Iconic Intercultural Food; Data from: Elizabeth Monroe Boggs: From Quantum Chemistry to the Manhattan Project

Additional selected datasets:
Simons Collaboration on Cracking the Glass Problem RDR Collection
Data from: Decorrelation of the static and dynamic length scales in hard-sphere glass formers (500th dataset published in RDR)

The RDR team thanks Dr. Charbonneau for his participation and supporting data publishing at Duke. For questions about the RDR contact: datamanagement@duke.edu

 

Announcing: Duke Research Data Repository 2.0

We are happy to announce that the Duke Research Data Repository (RDR) has updated our platform to provide enhancements for data depositors. The platform is implemented in partnership with Duke University Libraries and TIND, a spinoff of CERN (see the press release). Below we explain what you need to know about this change.

We also have this short video demonstrating key changes and features.

What is changing?

  • Homepage: There is a new landing page with slightly different navigation. All key documentation can be found under the “About” or “Resources” pages. You can still reach our site at research.repository.duke.edu.
  • Links: While URLs have changed for datasets, all DOIs assigned to datasets will continue to work. We always encourage you to use DOIs (vs. URLs from your browser window) as those are the most stable and persistent!
  • Persistent Identifiers (PIDs): All datasets will continue to receive DOIs, but we will also provide more robust support for ORCIDs (for unique identification of people) and RORs (for unique identification of organizations/institutions) for better tracking of research outputs and to comply with upcoming PID funder requirements
  • Metadata: We will now integrate more DataCite metadata. Key metadata (descriptive information) will remain primarily the same; however, the form for describing your data will now have more built-in features for metadata standardization and compliance with best practices (see PIDs above).
  • Dataset structure: All datasets now have a single landing page (no-sub pages as with the current platform). This resulted in certain datasets being slightly remodeled (e.g., folders zipped) to accommodate the new dataset structure. However, no files have been changed and data integrity, reuse, and reproducibility were prioritized in this process.

What do depositors need to know?

  • File upload: For smaller datasets, you will now upload your files directly within the web form prior to hitting the “Submit” button (vs. via a Box link in the previous workflow). For larger datasets, upload will continue to be facilitated via Globus.
  • Organization: Datasets will either need to be flat (no folders) or folders (if important to retain for access/reproducibility) will need to be packaged (e.g., zipped/tarred/etc) prior to upload.
  • ORCIDs: To allow linking ORCIDs with datasets, we strongly encourage all dataset authors to have an ORCID prior to beginning a data submission. Go to the ORCID website to get your ORCID today!

What is new and improved?

  • Submission dashboard: A new submission dashboard will allow depositors to track all current and past deposits in one easy location.
  • Versioning: A new versioning module will allow depositors to more easily request the creation of a new version of their dataset. All previous versions will still be retained and the new version will receive a new DOI.
  • File Previews: Previewing common file formats (e.g., tabular files, images, PDFs, zips) will now be available on the dataset landing page.
  • Data Citations: You will now be able to copy data citations in a wide variety of citation styles.

What is staying the same?

  • Curation by RDR staff: Datasets will continue to be reviewed by RDR staff prior to publication to help researchers make their data as FAIR (Findable, Accessible, Interoperable, and Reusable) as possible.
  • Preservation and retention: Data will continue to comply with the DUL preservation policy including redundant copies, fixity checks, and over all information security as required by Duke. Our stated retention policy also will not change.
  • Globus download: For datasets over a certain size threshold, users will continue to have the option to download those files via Globus (in addition to the option to download over the web browser). If the “Download from Globus” button appears at the top of a dataset, we encourage using Globus as downloading large scale data over a web browser has some challenges (e.g., timeouts, etc.).
  • Embargoes: Depositors can continue to request embargoes for up to one year and we can continue to facilitate access to embargoed files for journal reviewers. Embargoed file names will now be viewable on the dataset page but cannot be downloaded until the embargo is lifted.
  • Collections: Project-specific collections will still be supported and can be requested by emailing datamanagement@duke.edu.

The RDR curation team is excited to bring you this new and improved system and look forward to continuing to support data sharing, curation, and reproducibility for Duke generated data.

Please don’t hesitate to reach out with any questions at datamanagement@duke.edu.

Helenmary Sheridan, Research Data Management Consultant

Helenmary Sheridan, research data management consultant
Need help curating your data or identifying a repository to share data? Helenmary can help!

CDVS welcomes Helenmary Sheridan as the third member of the research data management (RDM) team. Helenmary joined Duke in August 2024 to help the library scale up classes, group trainings, and individualized consultations on RDM topics including NIH data management plans, data sharing in repositories such as the Duke Research Data Repository, and improving research reproducibility through documentation. Her position is supported by the Compute and Data Services Alliance for Research (CDSA), a new cross-campus initiative to support researchers with their computational needs.

Prior to joining Duke, Helenmary was the Data Services Librarian at the health sciences library at the University of Pittsburgh, where she provided data management training to faculty, staff, and students across the health disciplines. She has nearly ten years of experience working with scientific metadata and file formats, especially for data from imaging research (biomedical and otherwise.)

Helenmary’s favorite part of her job is teaching, especially Introduction to Research Data Management workshops for new graduate students and faculty that may be their first formal experience with research data methods. “It sounds like a dry subject,” she says, “so I love to see how excited researchers get when they realize how much easier these tools can make their lives.” You can contact Helenmary through the CDVS inbox at: askdata@duke.edu.

Local Data Repository Infrastructure Matters

Over the past seven years, the Duke University Libraries Research Data Curation Program and the Duke Research Data Repository (RDR) have developed into essential resources for the Duke research community. What began as a “proof of concept” turned into a robust, well-used data publishing service meeting both publisher and funder requirements and furthering Duke’s commitment to a “culture of open science and open scholarship to ensure transparency and accountability in research.”

The Duke Research Data Repository (RDR) has published over 310 datasets in a myriad of scientific disciplines including chemistry, biology, biomedical engineering, marine science, and medicine. Additionally, the RDR hosts datasets associated with articles published in PLOS, Nature, and PNAS and funded by NIH and NSF, their multiple sub-agencies and institutes, and others. The RDR provides a DOI for all datasets, and commits to long-term access and retention of data and curates all data based on the Data Curation Network (DCN) CURATED model.  As members of the renowned Data Curation Network (DCN), we leverage the expertise of a multi-institutional consortium that shares expertise to increase the quality of data curation across all DCN affiliated repositories.  The Duke Libraries are proud to be able to offer this local resource that supports the needs of Duke researchers who are not sufficiently served by disciplinary, data type or funder-based resources. In addition to providing the platform, we also provide front-line services for data management planning, data curation, disclosure risk review and referrals .

As the culture and landscape of data sharing evolves, researchers have many different repository options – from funder-sponsored repositories to discipline/community specific repositories to generalist repositories. Occasionally, journal publishers and funding agency Program Officers have questioned the suitability of institutional data repositories for long-term data sharing and preservation. To communicate the value of institutional resources focused on data sharing, we and other members of the DCN collaboratively wrote a letter to Science arguing that institutional data repositories provide valuable local infrastructure for researchers needing to meet data publishing guidelines.

In this letter, we detail how our repositories align with the FAIR guiding principles and commit to providing sustainable access to data critical for research reproducibility. This letter was published in the September 13 Issue of Science – Institutional Data Repositories are Vital (DOI: 10.1126/science.adr0789, open access copy available at: https://hdl.handle.net/11299/265639). The DCN has also published research that examined what researchers valued about their institutional data repositories and the services they provide. As one researcher noted

“I am thankful and excited for the help in curation…I see that teamwork in this final step of research means that the best possible version of the material will be available to future generations..”

All this to say, should you, as a researcher, need to demonstrate that an institutional data repository is an acceptable strategy for sharing data, we encourage you to cite the Science letter and reference the Duke Research Data Repository’s documentation clarifying how the RDR approaches compliance with the NIH Desirable Characteristics for Data Repositories . Comments or questions about the RDR can be sent to datamanagement@duke.edu.

Publications referenced:

Jen Darragh et al. (2024). Institutional data repositories are vital. Science 385,1174-1174(2024). DOI:10.1126/science.adr0789

Marsolek W, Wright SJ, Luong H, Braxton SM, Carlson J, Lafferty-Hess S (2023). Understanding the value of curation: A survey of researcher perspectives of data curation services from six US institutions. PLoS ONE 18 (11): e0293534. https://doi.org/10.1371/journal.pone.0293534

The Duke Research Data Repository Celebrates its 200th Data Deposit!

The Curation Team for the Duke Research Data Repository is happy to present an interview with Dr. Thomas Struhsaker, Retired Adjunct Professor of Evolutionary Anthropology.

CC-BY Thomas Struhsaker, Medium Juvenile Eating Charcoal, July 1994, Jozani

Dr. Struhsaker’s dataset, Digitized tape recordings of Red colobus and other African forest monkey species vocalizations, was the 200th dataset to be added to the Duke Research Data Repository. I worked closely with Tom to arrange and describe this collection. He hopes to be adding even more in the near future as he winds down his career. Tom might not know this, but his dataset has been tweeted about 36 times at this point and has been viewed 336 times since August. Ever the humble scientist, I did not know until I saw the tweets that Tom was the winner of the 2022 President’s Award from the American Society of Primatologists (congratulations Tom!).

I started my interview with Dr. Struhsaker as one typically would – by asking him to tell me about himself and his field of research. He laughs and says “Oh boy, where to begin? You’re talking half a century here.” I could listen to Tom talk for hours about his experiences as a young field biologist at a time when primatology was just figuring itself out. Tom went about his work as a naturalist – do not interfere, observe and learn. He spent 25 years in Africa (spanning 56 years from 1962-2018), observing many different species of animals, not just primates. For 18 of these 25 years Tom lived in Uganda as a full-time resident, including during the reign of Idi Amin, one of the most brutal rulers in modern history. Idi Amin aside, Tom thought that the Ugandans were some of the best folks to work with regarding conservation in Africa due to their dedication to higher education (Makerere University) with growing generations of students and the establishment of Kibale National Park. I cannot do Tom’s fascinating life justice in just this short blog post, so I encourage you to read Tom’s 2022 article, The life of a naturalist (full text access available through NetID login) and his memoir, I remember Africa: A field biologist’s half-century perspective (Perkins & Bostock Library – Duke Authors Display – QH31.S79 A3 2021). What I can tell you, at least from my perspective, is that Tom has led a life passionate about nature, wanting to know everything he could from our cohabiters on this planet and how we can best live together.  If you would like your own copy it can be purchased here.

Tom recorded these vocalizations between 1969-1992. He thought it was really important to do so because they are key to understanding communication and the social life of primates. Analysis of these recordings led Tom to conclude that among African monkeys vocalizations are relatively stable characters from an evolutionary perspective and, therefore, important in understanding phylogenetic relationships.  As for archiving and sharing the recordings of these vocalizations, Tom didn’t initially have that in mind. He instead followed the more traditional academic route of publishing articles including spectrograms, and his conclusions about the meaning of the vocalizations. Over the last two years as Tom began thinking about the legacy of his materials, he realized that while the visual representations are useful to share for analysis, it is just not the same as listening to the sounds themselves. Why not archive them to make it possible for others to hear them?

“He realized that while the visual representations are useful to share for analysis, it is just not the same as listening to the sounds themselves. Why not archive them to make it possible for others to hear them?”

With increasing human populations, deforestation, climate changes, etc., some of these animals (like the Red Colobus) have become critically endangered, and these recordings might be the only way future generations will ever be able to hear these animals. Tom’s recordings were made using reel to reel tapes on very large and heavy tape recorders with 12 D-Cell batteries. Crawling through the forest with these machines in addition to a large boom microphone was no easy feat. With the help of the Macaulay Library (Mr. Matthew Medler in particular), several of the original tapes were digitized to the high-quality WAV files we have in the collection. Tom has also augmented the collection with his own MP3 recordings. He hopes to have more WAV format from Macaulay Library in the future.

Tom did not initially know where to archive these vocalizations as they weren’t in scope for MorphoSource (another Duke-based repository for 3D imaging) where Tom will soon have a collection of red colobus monkey images available. Thanks to a suggestion from his neighbor Ben Donnelly, he reached out to the Duke Research Data Repository Curation Team (thanks for being a great colleague Ben!). This is where I (Jen Darragh), the author, come in.

Tom and I worked together over the course of a couple months to build his data deposit. Perhaps somewhat self-servingly, I asked him how he found the process. He stoked my ego with both a “fantastic, and easy peasy.” He said he would recommend us to anyone as we do our best to make the process as clear and pain-free as possible. Aw shucks Tom. You are one of my favorite depositors to work with, too.

I asked Tom what would he advise for early career researchers and those just getting started in the field when it comes to data sharing and archiving. He said that he is seeing increasing requirements as part of publishing (he’s right) and he’s in favor, as long as the person who collected the data is credited (cite properly!) and consulted when possible (collaboration is good). It’s important to advance the sciences. Repositories help to encourage good citation practices in addition to the preservation of important data for the long-term.

CC-BY Thomas Struhsaker. Medium-large juvenile red colobus (eating bark of bottle brush tree, Kanyawara, Kibale National Park, Uganda.

Tom also mentioned some longitudinal data he had collaborative built over the years with colleagues and that continues to be built upon. His experience of archiving his vocalization recordings with us (and his images with MorphoSource) got him thinking that repositories are a wonderful option to ensure that these important materials continue to persist and be used. He has thought of at least three important datasets and plans to reach out to his collaborators about archiving these data either with us in the Duke RDR, or in another formal repository of their choosing.

Tom recently shared with me a collection of photographs that he has taken in the same spot in Kibale from 1976-2018 that shows how the area went from bare grassland to a low stature forest (pre-conservation to post-conservation efforts). He has shared these with his colleagues directly to show the fascinating change over time. He now hopes to share them more broadly through the Duke RDR (forthcoming, we have some processing to do). Perhaps someone will be inspired to animate the images and then share back with us.

To close the interview, I asked Tom what his favorite animal was. I think it’s no surprise that he likes them all; there are so many he likes for different reasons, some subtle, some not (“some insects are damn weird”) and some just do incredibly interesting things. The diversity is what he loves.

Struhsaker, T. T. (2022). Digitized tape recordings of Red colobus and other African forest monkey species vocalizations. Duke Research Data Repository. https://doi.org/10.7924/r4pv6nm9f

 

Dr. Mark Palmeri: An honest assessment of openness

This post is part of the Duke Research Data Curation Team’s ‘Researcher Highlight’ series.

In the field of engineering, a key driving motivator is theDr. Mark Palmeri urge to solve problems and provide tools to the community to address those problems. For Dr. Mark Palmeri, Professor in Biomedical Engineering at Duke University, open research practices support the ultimate goals of this work, and helps get the data into the hands of those solving problems: “It’s one thing to get a publication out there and see it get cited. It’s totally another thing to see people you have no direct professional connection to accessing the data and see it impacting something they’re doing…”

Dr. Palmeri’s research focuses on medical ultrasonic imaging, specifically using acoustic radiation force imaging to characterize the stiffness of tissues. His code and data allow other researchers to calibrate and validate processing protocols, and facilitate training of deep learning algorithms. He recently sat down with the Duke Research Data Repository Curation Team to discuss his thoughts on open science and data publishing.

“It’s one thing to get a publication out there and see it get cited. It’s totally another thing to see people you have no direct professional connection to accessing the data and see it impacting something they’re doing…”

With the new NIH data management and sharing policy on the horizon, many researchers are now considering what sharing data looks like for their own work. Palmeri highlighted some common challenges that many researchers will face, such as the inability to share proprietary data when working with industry partners, de-identifying data for public use (and who actually signs off on this process), the growing scope and scale of data in the digital age, and investing the necessary time to prepare data for public consumption. However, two of his biggest challenges relate to the changing pace of technology and the lack of data standards.

ClockWhen publishing a dataset, you necessarily have a static version of the dataset established in space and time via a persistent identifier (i.e., DOI); however, Palmeri’s code and software outputs are constantly evolving as are the underlying computational environments. This mismatch can result in datasets becoming out of sync with the coding tools, thereby affecting future reuse and ultimately keeping things up-to-date takes time and effort. As Palmeri notes, in the fast-paced culture of academia “no one has time to keep old project data up to snuff.”

Likewise, while certain types of data in medical imaging have standardized formats (e.g., DICOM), for the images Palmeri is creating from raw signal data there are no ubiquitous standards. This creates problems for data reuse. Palmeri remarks that “There’s no data model that exists to say what metadata should be provided, in what units, what major fields and subfields, so that becomes a major strain on the ability to meaningfully share the data, because if someone can’t open it up and know how to parse it and unwrap it and categorize it, you’re sharing gigabytes of bits that don’t really help anyone.” Currently, Dr. Palmeri is working with the Quantitative Imaging Biomarkers Alliance and the International Electrotechnical Commision (IEC) TC87 (Ultrasonics) WG9 (Shear Wave Elastography) to create a public standard for this technology for clinical use.

Ultrasound scanner images
Image processing example using MimickNet

Regardless of these challenges, Palmeri sees many benefits to publicly sharing data including enhancing “our internal rigor even just that little bit more” as well as opening “new doors of opportunity for new research questions…and then the scope and impact of the work can be augmented.” Dr. Palmeri appreciates the infrastructure provided by the Duke University Libraries to host his data in a centralized and distributed network as well as the ability to cite his data via the DOI. As he notes “you don’t want to just put up something on Box as those services can change year to year and don’t provide a really good preserved resource.” Beyond the infrastructure, he appreciates how the curation team provides “an objective third party [to] look at things and evaluate how shareable is this.”

“you don’t want to just put up something on Box as those services can change year to year and don’t provide a really good preserved resource.”

Within the Duke Research Data Repository, we have a mission to help Duke researchers make their data accessible to enable reproducibility and reuse. Working with researchers, like Dr. Palmeri, to realize a future where open research practices lead to a greater impact for researchers and democratizes knowledge is a core driving motivator. Contact us (datamanagement@duke.edu) with any questions you might have about starting your own data sharing adventure!

Share More Data in the Duke Research Data Repository!

We are happy to announce expanded features for the public sharing of large scale data in the Duke Research Data Repository! The importance of open science for the public good is more relevant than ever and scientific research is increasingly happening at scale. Relatedly, journals and funding agencies are requiring researchers to share the data produced during the course of their research (for instance see the newly released NIH Data Management and Sharing Policy). In response to this growing and evolving data sharing landscape, the Duke Research Data Repository team has partnered with Research Computing and OIT to integrate the Globus file transfer system to streamline the public sharing of large scale data generated at Duke. The new RDR features include:

  • A streamlined workflow for depositing large scale data to the repository
  • An integrated process for downloading large scale data (datasets over 2GB) from the repository
  • New options for exporting smaller datasets directly through your browser
  • New support for describing and using collections to highlight groups of datasets generated by a project or group (see this example)
  • Additional free storage (up to 100 GB per deposit) to the Duke community during 2021!

While using Globus for both upload and download requires a few configuration steps by end users, we have strived to simplify this process with new user documentation and video walk-throughs. This is the perfect time to share those large(r) datasets (although smaller datasets are also welcome!).

Contact us today with questions or get started with a deposit!

Publish Your Data: Researcher Highlight

This post was authored by Shadae Gatlin, DUL Repository Services Analyst and member of the Research Data Curation Team.

Collaborating for openness

The Duke University Libraries’ Research Data Curation team has the privilege to collaborate with exceptional researchers and scholars who are advancing their fields through open data sharing in the Duke Research Data Repository (RDR). One such researcher, Martin Fischer, Ph.D., Associate Research Professor in the Departments of Chemistry and Physics, recently discussed his thoughts on open data sharing with us. A trained physicist, Dr. Fischer describes himself as an “optics person” his work ranges from developing microscopes that can examine melanin in tissues to looking at pigment distribution in artwork. He has published data in the RDR on more than one occasion and says of the data deposit process that, “I can only say, it was a breeze.”

“I can only say, it was a breeze.”

Dr. Fischer recalls his first time working with the team as being “much easier than I thought it was going to be.” When Dr. Fischer and colleagues experienced obstacles trying to setup OMERO, a server to host their project data, they turned to the Duke Research Data Repository as a possible solution to storing the data. This was Dr. Fischer’s first foray into open data publishing, and he characterizes the team as being  responsive and easy to work with. Due to the large size of the data, the team even offered to pick up the hard drive from Fischer’s office. After they acquired the data, the team curated, archived, and then published it, resulting in Fischer’s first dataset in the RDR.

Why share data?

When asked why he believes open data sharing is important, Dr. Fischer says that “sharing data creates an opportunity for others to help develop things with you.” For example, after sharing his latest dataset  which evaluates the efficacy of masks to reduce the transmission of respiratory droplets, Fischer received requests for a non-proprietary option for data analysis instead of using the team’s data analysis scripts written for the commercial program Mathematica. Peers offered to help develop a Python script, which is now openly available, and for which the developers used the RDR data as a reference. As of January 2021, the dataset has had 991 page views.

Dr. Fischer appreciates the opportunity for research development that open data sharing creates, saying, “Maybe somebody else will develop a routine, or develop something that is better, easier than what we have”. Datasets deposited in the RDR are made publicly available for download and receive a permanent DOI link, which makes the data even more accessible.

“Maybe somebody else will develop a routine, or develop something that is better, easier than what we have.”

In addition to the benefits of long-term preservation and access that publishing data in the RDR provides, Dr. Fischer finds that sharing his data openly encourages a sense of accountability. “I don’t have a problem with other people going in and trying, and making sure it’s actually right. I welcome the opportunity for feedback”. With many research funding agencies introducing policies for research data management and data sharing practices, the RDR is a great option for Duke researchers. Every dataset that is accepted into the RDR is carefully curated to meet FAIR guidelines and optimized for future reuse.

Collaborating with researchers like Dr. Martin Fischer is one of the highlights of working on the Research Data Curation team. We look forward to seeing what fascinating data 2021 will bring to the RDR and working with more Duke researchers to share their data with the world.

Dr. Fischer’s Work in the Duke Research Data Repository:

  • Wilson, J. W., Degan, S., Gainey, C. S., Mitropoulos, T., Simpson, M. J., Zhang, J. Y., & Warren, W. S. (2019). Data from: In vivo pump-probe and multiphoton fluorescence microscopy of melanoma and pigmented lesions in a mouse model. Duke Digital Repository. https://doi.org/10.7924/r4cc0zp95
  • Fischer, E., Fischer, M., Grass, D., Henrion, I., Warren, W., Westman, E. (2020). Video data files from: Low-cost measurement of facemask efficacy for filtering expelled droplets during speech. Duke Research Data Repository. V2 https://doi.org/10.7924/r4ww7dx6q

Got Data? Data Publishing Services at Duke Continue During COVID-19

While the library may be physically closed, the Duke Research Data Repository (RDR) is open and accepting data deposits. If you have a data sharing requirement you need to meet for a journal publisher or funding agency we’ve got you covered. If you have COVID-19 data that can be openly shared, we can help make these vital research materials available to the public and the research community today. Or if you have data that needs to be under access restrictions, we can connect you to partner disciplinary repositories that support clinical trials data, social science data, or qualitative data.

Speaking of the RDR, we just completed a refresh on the platform and added several features!

In-line with data sharing standards, we also assign a digital object identifier (DOI) to all datasets, provide structured metadata for discovery, curate data to further enhance datasets for reuse and reproducibility, provide safe archival storage, and a standardized citation for proper acknowledgement.

Openness supports the acceleration of science and the generation of knowledge. Within the libraries we look forward to partnering with Duke researchers to disseminate their research data! Visit https://research.repository.duke.edu/ to learn more or contact datamanagement@duke.edu with any questions.