Patrick Charbonneau on why he shares his data with the world

This story is part of a fall 2026 series celebrating 500 published datasets in the Duke Research Data Repository by spotlighting 5 Duke researchers who make their work more transparent, discoverable, and reusable.

Dr. Patrick Charbonneau maintains appointments in both Chemistry and Physics. Dr. Charbonneau published the 500th dataset within the Duke Research Data Repository (RDR) and is the top depositor. He shared with us his thoughts on data sharing and the use of the RDR.

Dr. Patrick CharbonneauTell us a little about the research and the kinds of data shared through the RDR. What should someone outside your field know about them?

My group works at the intersection of theoretical chemistry and physics, using statistical mechanics and computational methods to understand complex systems. The datasets we share typically contain the numerical data behind published figures, along with code and documentation needed to understand or reproduce the results.

Why do you choose to share your research data? Has your thinking about data sharing changed over time?

Early in my career, archiving meant asking students and postdocs, usually as they were leaving, to organize their files on a departmental drive. I eventually realized that this approach was both stressful and fragile. We therefore shifted to preparing data, code, and documentation alongside each manuscript, so that sharing becomes part of publishing rather than something done retrospectively. An unexpected benefit is that preparing a dataset for others to use provides an additional check on the research itself.

What do you hope other researchers, or other audiences, might be able to do with the data you’ve shared?

Dr. Charbonneau with then graduate student, Lin Fu, in 2012

I want our publications to remain reusable without having to track down a former student, postdoc, or me(!) years later. Surprisingly, this happened twice just this summer for papers from before we routinely deposited datasets; I have since deposited those data. But I also like that we cannot entirely predict what the data might eventually be useful for. A dataset created to answer one question may later become useful for comparison, methodological development, teaching, or a question we had never considered.

What has your experience been like working with the Duke Research Data Repository team? Is there anything about the publishing or curation process that has been particularly useful to you?

The RDR team has made data publication remarkably easy to incorporate into our research workflow. The curatorial oversight is particularly valuable: having another set of eyes on the organization, documentation, and long-term usability of a dataset helps turn files associated with a paper into a durable research product. Over time, working with the RDR has made depositing data simply part of what we consider completing a publication.

What would you say to another Duke researcher who is considering sharing their data?

Start small. There is no need to solve every possible question about what should be preserved before making a first deposit. Sharing even the numerical values behind published figures is a useful beginning, and once the process becomes part of the publication workflow, expanding what you share becomes much easier. In my experience, the cost is small, especially with the support of the RDR team, while the benefits for transparency, preservation, and future reuse can be substantial. Also, LLMs now make writing useful README files easier than ever.

A tray of pieces of sucre à la crème
A tray of pieces of sucre à la crème, the data from the associated article is published in the RDR

Has anything interesting or unexpected happened as a result of sharing data?

In recent years, I started publishing articles on the history of science and related questions. Surprisingly, the datasets I deposited to accompany those articles—oral history interviews, textual corpora, and archival documents—are among the most consulted of all my deposits. I certainly didn’t expect my humanities datasets to be the popular ones!

Charbonneau’s humanities datasets: Data from: Sucre à la crème: Origin and Trajectory of an Authentic Québec Confection; Data from: Pralines des Voyageurs: An Iconic Intercultural Food; Data from: Elizabeth Monroe Boggs: From Quantum Chemistry to the Manhattan Project

Additional selected datasets:
Simons Collaboration on Cracking the Glass Problem RDR Collection
Data from: Decorrelation of the static and dynamic length scales in hard-sphere glass formers (500th dataset published in RDR)

The RDR team thanks Dr. Charbonneau for his participation and supporting data publishing at Duke. For questions about the RDR contact: datamanagement@duke.edu