Contents
From signposts to data maps
Paul frames the same problem from the systems side. A TMF has never been purely a document repository. It has always pointed to information that lives in other systems: CTMS, safety databases, monitoring platforms.
"Traditionally, in our trial master file systems, we've created what we call signposts, which say, okay, this record exists and it exists in this system, and this is what it's called, and this is what it's documenting," Paul said. "That's all well and good all the time that those systems are up and running and live. But as soon as you decommission those systems, all of a sudden, your signpost no longer points anywhere."
Paul connects this directly to a shift the industry is only starting to reckon with: essential-records work is more than an inventory. It functions as a data map.
Understanding where information lives and how it flows between systems, a topic that came up explicitly at a recent MHRA, FDA, and Health Canada symposium in Ottawa, is becoming as important as knowing what the record is. "We're moving away from the document paradigm," Paul said. "We're moving towards a data paradigm."
More and more of what a trial produces is a dataset rather than a document. The standard doing the most to formalize that shift is USDM, the Unified Study Definition Model, now in its fourth iteration. USDM was built around digital protocol, but its logical data model has grown well past that and is expected to eventually incorporate TMF-specific elements.
Standardized metadata, Paul argues, is the ingredient that makes any of this interoperable rather than merely connected: "One of the keys to successful interoperability is standard metadata. [Without] standard metadata, it's not going to happen."
I think that distinction, interoperability versus integration, is worth sitting with. Integration is a point-to-point connector between two systems. Interoperability means every system in the ecosystem speaks a common, standards-based language by default. One scales. The other doesn't.
Alan made the same point with a household analogy:
"That's like when you renovate your kitchen and you buy all your white goods from the same supplier, and you get a really good fridge and a really good stove, but they're not very good at dishwashers and they don't understand microwave ovens, and you end up with something that's subpar. […] You prefer to be able to pick what suits your study best, and the way to make it work when you have many different systems is interoperability."
The cost of a silo is not always measured in hours
I wanted Rachana's perspective specifically because it comes from patient access rather than systems architecture, and it reframes what's at stake.
Clinical operations teams tend to describe siloed data in terms of inefficiency and inspection readiness. Those costs are real. But Rachana named a cost that 'rarely shows up' in a business case: the patient who never gets found.
She's seen this pattern consistently in rare disease and underserved communities, where patient registries don't connect to trial matching platforms and site data sits disconnected from sponsor systems. Community health settings often serve a higher proportion of patients of color, yet trial matching data reaches them last, if at all.
"The cost of siloed data isn't just the slower timelines for sponsors," Rachana said. "It is that diversity gap, and that compounds with every disconnected system."
The result isn't only a slower timeline. It's a narrower evidence base for the therapies coming out of that research, built on the population the system happened to reach rather than the population the therapy is meant to serve. In her words, the evidence base "reflects again back on that narrow slice of people."
Her point about trust is the one I keep coming back to. Connecting smaller sites and patient advocacy organizations into the ecosystem as full nodes, rather than treating them as an afterthought once the sponsor-CRO-vendor pipeline is built, is what makes patients willing to enter that pipeline in the first place.
"Making it easier for smaller sites, smaller community health organizations, patient advocacy groups to be part of that network as interconnected nodes, and not just an afterthought, is going to build on that trust, is going to generate that support system, and ultimately a connected ecosystem," Rachana said.
Ownership does not disappear when work gets delegated
No matter how work gets distributed across CROs and vendors, GCP places the accountability with the sponsor. A sponsor is able to delegate the work. The responsibility stays put, and the only place that delegation gets formalized is in the contract.
"However you want to slice and dice it, at the end of the day, it's the sponsor who's responsible," Alan said. "And the only way that they can delegate that responsibility is in their contracts. They have to ensure that they know exactly what's going on in their study."
Alan's group made this one of eleven positions in a paper it published roughly five years ago, and the essential records checklist exists partly to give that principle teeth: a sponsor who names exactly what must be extracted from every system in use is able to hold every CRO and vendor to that specification in writing.
Paul draws a parallel to data management, where data transfer agreements spell out exactly what moves between organizations and when. TMF custody stands to benefit from the same discipline: an agreement stating, for example, that monitoring reports stay inside the CTMS along with their metadata and audit trail until study close, at which point they transfer to the sponsor's system of record as the authoritative source. Without that clarity, an inspector asking which copy is authoritative doesn't have a clean answer.
ICH E6(R3) places oversight responsibility on the sponsor. Study Oversight in eTMF Connect gives you a documented record of that oversight for each study. See the workflow.
Audit trails are the harder problem, and PDFs are not the answer
Regulatory guidance is specific about what an audit trail must capture — creation, modification, and deletion of records, per 21 CFR Part 11 and Annex 11. It's far less specific about what happens to that audit trail once records move, archive, or transfer custody during a system decommission.
Dumping an audit trail into a spreadsheet doesn't guarantee its integrity, and it doesn't make the record easy to interrogate later if a question comes up during inspection.
Alan made one point on this that I think deserves repeating: regulators have told his group directly, in meetings on both sides of the Atlantic, that they don't want PDFs and won't flag their absence in an inspection finding.
"I've talked to the regulators both in Europe and in North America," he said. "They don't want PDFs. They don't expect PDFs. You won't get an audit finding if you don't have PDFs. The PDFs take an enormous amount of space, and everything else, your audit trails, everything else compared to your PDFs is negligible."
"Sponsors keep asking vendors for them anyway, out of habit more than requirement. The volume problem in TMF archives is rarely the audit trail. The PDFs are the problem."
Where regulators are already heading
Data flow diagrams are now showing up as an expected part of regulatory submissions, a sign that authorities are already thinking in terms of connected ecosystems rather than isolated document sets.
"It's become visible in the regulations and their requirements for data flow diagrams," Alan noted. "You have to have that as part of your information that you put forward to them now."
Paul expects that trajectory to continue: "Having better connected information allows us to better reconcile that information, identify issues before inspections occur. We can perform a lot of analytics on data to detect anomalies and things that just don't quite make sense, prior to inspections and in our day to day operations."
Rachana put the destination in the clearest terms of our discussion: "At some point, we should stop talking about interoperability," she said. "It should be a point where the systems are connected. We have a system with the patient at the center of the ecosystem." I agreed. She went further: "Data is following the study, not the other way around."
None of this arrives through a single vendor connecting its own suite of systems end to end. Interoperability comes from standards that let best-of-breed systems, chosen for what they each do well, exchange clinical data and essential records in a common format.
This is the direction the CDISC TMF Standard Model, USDM, and the essential records work are converging on. It's the standard we build eTMF Connect against.