Est.

Funding Agency Reproducibility Mandates Beyond NIH

Federal agencies now require researchers to share data and methods at publication, not months later.

Contributing Editor · · 10 min read
Cover illustration for “Funding Agency Reproducibility Mandates Beyond NIH”
Reproducibility Policy · September 30, 2026 · 10 min read · 2,149 words

A single directive from the Office of Science and Technology Policy set this entire structure in motion. In August 2022, OSTP issued a memorandum titled "Ensuring Free, Immediate, and Equitable Access to Federally Funded Research," instructing every federal agency with any research and development expenditures to rewrite its public-access and data-sharing policies so that publications and the data behind them would be freely available at the time of publication, with no embargo period. Agencies did not move at the same pace. Policies were expected to be finalized by the end of 2024 and put into practice by the end of 2025. The obligations researchers are now encountering are already in force. That timeline matters for anyone assuming they have a grace period left to sort out compliance.

The scale of what followed has only grown. The OMB's May 29, 2026 proposed rule to revise 2 CFR Part 200, the Uniform Guidance governing all federal grants, involves virtually every federal grantmaking agency; attorneys at Hogan Lovells characterized it as the most significant revision to federal grants policy in a generation. NIH's Rigor and Reproducibility framework, long treated by biomedical researchers as the benchmark for what funders expect, now sits as one expression of a much broader federal posture rather than the outer limit of what any given grantee must satisfy. Researchers funded by NSF, DOE, DOD, or a dozen other agencies now face simultaneously active, overlapping mandates that vary by agency.

What "reproducibility" means across these mandates

The word "reproducibility" does not mean the same thing to every agency writing it into policy, and that inconsistency is not a semantic footnote. Some frameworks define it narrowly, as the ability to take the same data and the same methods and arrive at the same result. Others extend it further, requiring independent replication using newly generated data. Barba (2018) documented this exact ambiguity across scientific disciplines, and the consequence is practical rather than philosophical: what a mandate actually requires a lab to share depends entirely on which definition the funding agency has adopted.

NSF's GFA frames scientific integrity explicitly around reproducibility, transparency, communicativeness of error and uncertainty, collaborative and interdisciplinary approaches, skepticism of findings and assumptions, falsifiability, unbiased peer review, acceptance of negative results, and freedom from conflicts of interest, a definition that removed the prior PAPPG's language about inclusive environments and now anchors compliance to methodological standards. Notably, this definition dropped language from the prior PAPPG about inclusive environments, replacing it with a framework anchored squarely in methodology. NIH took a more granular route. Its April 2026 guidance breaks reproducibility into four application-level requirements: rigor of prior research, scientific rigor, attention to relevant biological variables, and authentication of key resources. Authentication of key resources requires critical materials to be validated for reliability: reagent identity, lot provenance, and quality-control data now count as required compliance documentation, not optional lab notebook detail. NIH consolidated all of this into a centralized Replication and Reproducibility Initiative hub, launched in early 2026, specifically to communicate these expectations across the biomedical research enterprise.

NSF's specific requirements under the GFA

NSF has undergone the most structurally significant policy overhaul of any agency outside NIH, and the timing is not theoretical. The old Proposal and Award Policies and Procedures Guide has been restructured into the Guidance on Financial Assistance, split into 26 separate guides, in what counts as the broadest reorganization NSF has made since 2014. Under PAPPG 24-1, Supplement 2, these changes apply to every award made on or after January 22, 2026.

The most consequential single change is the elimination of the 12-month publication embargo. Investigators must now deposit accepted author manuscripts and the underlying datasets in NSF's Public Access Repository at or before publication, not within a year afterward as the old rule allowed. NSF's Biological Sciences Directorate has been explicit that this requirement exists in service of reproducibility and transparency, tying the timing rule directly to what the agency calls Gold Standard Science. A new Research.gov tool for creating Data Management and Sharing Plans launched April 27, 2026; investigators must now use it to ensure data supporting NSF-funded publications are shared at the time of publication.

One part of the new guidance has not been fully worked out. The GFA mandates immediate open access to publications while simultaneously proposing to disallow the use of grant funds for "publication costs," a term the draft never defines. Guide 4 still lists reports, reprints, page charges, and illustrations as allowable expenses, leaving a direct contradiction unresolved as of the August 24 public comment deadline. Researchers budgeting for open-access fees should treat this as an open question to track until NSF clarifies which costs qualify.

NSF also widened the scope of its Responsible Conduct of Research training under Important Notice No. 149, adding research security awareness, export controls, and disclosure and reporting requirements to a training regime that previously had no national security component at all. For bench scientists, the sum of these changes is straightforward even if the policy language is not: reagent identity, lot data, and the protocols underlying a published result now have to be documented and made shareable at the time of submission, not assembled retroactively after a journal asks for supplementary materials.

How DARPA, DOE, and other agencies apply the same directive

Chapman University's library guide, tracking this landscape through 2026, lists active requirements at these agencies and offices: AHRQ, ASPR, CDC, DHHS, DOD, DOE's Education office, DOE's Energy office, DOT, EPA, FDA, IMLS, NASA, NCAR, NEH, NIH, NIST, NOAA, NSF, the Smithsonian Institution, USDA, USGS, and VA, confirming the mandate landscape is as wide as the OSTP memo intended. That breadth confirms the OSTP memo achieved exactly the reach it was written for.

Some agencies have built out frameworks that mirror NIH almost directly. ASPR requires investigators to submit an electronic version of the Author Accepted Manuscript to PubMed Central at the moment of acceptance, and underlying data must be freely available in public repositories, in machine-readable formats, at the time of initial publication. AHRQ takes a slightly different approach, requiring a formal data management plan as part of the grant application itself, covering primary data, samples, physical collections, and supporting materials. ARPA-H has gone the furthest of any agency in naming reproducibility as a mission, not just a policy line item. Its Intelligent Generator of Research program, launched May 5, 2026, describes a systemic effort to deliver gold-standard biomedical science faster, built around an AI-powered research ecosystem meant to accelerate breakthroughs. The common thread running through all of these frameworks is timing: data management and sharing plans are now required at the application stage, before any award is made.

What researchers must document about experimental inputs

Policy language eventually becomes a lab requirement, and this is where it lands. Across NSF, NIH, AHRQ, and ASPR, the mandates reach past the manuscript and into the bench itself: the reagents, protocols, materials, and lot identifiers behind a published result now have to be documented and shareable, not summarized in a paragraph of methods text. NIH's authentication of key resources requirement makes this explicit for biological materials, demanding confirmed identity and validation for anything like a cell line or antibody whose variability could shift results. NSF's DMSP requirement, now filed through Research.gov, covers data, software, samples, and other research outputs, language broad enough to directly implicate the physical reagents and reaction conditions used to generate a published dataset.

The failure mode is specific and increasingly common. A researcher who builds a paper on a commercial reagent kit with no published formulation, no lot-level quality-control data, and no deposited protocol is now sitting on a compliance gap under several mandates at once, for a simple reason: an independent lab cannot reproduce the work without that information. A catalog number dropped into a methods section used to be considered adequate description. The requirement now is machine-readable data deposited in a public repository at the time of publication, so a catalog number or a paragraph of prose referencing a proprietary formulation no longer satisfies a data-sharing plan.

Why reagent transparency is now compliance infrastructure

Lot-to-lot variability has quietly been the largest reproducibility threat sitting inside protein synthesis workflows for years. Most cell-free protein synthesis (CFPS) studies report yields well below what optimized systems can achieve, and manufacturer inconsistency in components such as purified tRNA has been flagged directly as a barrier to scaling these platforms. What used to be a scientific frustration is now a documentation gap that a funding agency can flag.

A study by Olsen and colleagues, published in Nature Communications, shows what compliance-ready reagent documentation looks like in practice. The team screened 1,231 different reagent combinations to arrive at a 12-component formulation, then demonstrated that it performed consistently across different batches of cell lysate, across different users and locations, and across the synthesis of more than 20 distinct proteins. That cross-lot, cross-lab robustness data is precisely the category of evidence NSF and ARPA-H mandates are now asking for. A simplified eCFPS system described in a bioRxiv study reduced essential reaction components from 35 to a core set of 7, demonstrating that formulation transparency and simplification are compatible with high performance.

What these mandates are asking for, in plain terms, is a system where a researcher at a different institution, working from a different reagent lot, can run the identical protocol and land on the identical result. That requires the formulation to be published, the lot-level QC data to be accessible, and the protocol itself to sit in a deposited repository rather than a lab notebook. OpenCFPS™ from Sepia Biosciences was built around that exact logic: published formulations, lot-level QC data made available to users, and pricing structured to make high-throughput protein synthesis accessible rather than gatekept behind proprietary black boxes. For labs anticipating a data-sharing plan audit, a reagent system built to that standard is a credible fit because its documentation was designed from the start to survive this kind of scrutiny.

Building a data plan that satisfies multiple agencies at once

The OSTP memo established one common floor across these agencies, so a data management and sharing plan built to satisfy NSF's requirements will overlap substantially with what AHRQ, ASPR, and NIH each expect. The goal for any multi-funded lab should be a single layered plan that covers every grant. A cross-agency-compliant DMSP needs to address several core elements:

  • Data types and formats: what data will be generated, in what machine-readable format, and which repository will hold it.
  • Timing: data and manuscripts deposited at or before publication, not after, the embargo-free standard now operative under NSF GFA and OSTP-aligned agency policies.
  • Reagent and materials documentation: lot identifiers, formulation sources, QC data, and protocols deposited in a named repository or supplementary package.
  • Authentication plans: a description of how critical reagents and resources were validated, following NIH's four-element framework.
  • Software and computational workflows: code and analysis scripts deposited with enough documentation that another lab could reproduce the analysis independently.

NSF's Research.gov DMSP tool, live since April 27, 2026, is the practical entry point for NSF-funded labs, and its structure translates reasonably well into templates for other agencies' plans. The Whole Tale project, described by Brinckman and colleagues, offers a conceptual model that is directly relevant: "living publications" that unite data products with research articles, executable objects integrating data and computational details, are the end state these mandates are pushing toward. For protein synthesis work specifically, a minimum viable compliance package includes the reaction formulation with lot numbers and sources for every component, the protocol as a deposited document, QC data for key reagents, and the raw output data, all linked to the publication at the moment of submission. Labs that build this infrastructure once, rather than reassembling it under deadline pressure for each new grant, end up positioned to satisfy whichever agency is funding the next project.

Where CFPS fits

Cell-free protein synthesis has a structural advantage that cell-based expression systems simply do not share. Because the reaction happens in an open, in vitro environment rather than inside a living cell, every component added to it can, in principle, be listed, lot-tracked, and deposited alongside the publication. Documenting the internal state of a living cell during expression is a far messier problem than documenting a defined mixture of reagents in a tube.

That advantage is not incidental for certain classes of proteins. Lot-to-lot variability in complex reagents is the single largest reproducibility threat in protein synthesis workflows: most CFPS studies report yields well below what optimized systems can achieve, and manufacturer variability in components like purified tRNA has been identified as a direct barrier to implementation at scale. It means the platform's transparency, every component visible, every lot traceable, gives labs a genuine head start on satisfying mandates that were written with exactly this kind of documentation in mind, provided the underlying reagent systems are built and published with that standard in view from the outset.

Sources

  1. NSF mandates immediate open access, disallows grant funds to pay for it
  2. Funding Agencies with Open Access and Data Sharing Requirements - Open Access and Data Sharing Mandates - LibGuides at Chapman University
  3. NIH Launches New Central Resource to Support Replication and Reproducibility | Grants & Funding
  4. Terminologies for Reproducible Research
  5. Embedding Replication and Reproducibility Throughout NIH Research: Key Reminders for Applications, Awards, and a New Highlighted Topic | Grants & Funding
  6. Computing Environments for Reproducibility: Capturing the "Whole Tale"
  7. Strengthening Replication and Reproducibility of NIH-funded Research | National Institutes of Health (NIH)
  8. Upcoming public access requirements for federally funded publications and data - University Library

More in Reproducibility Policy