In the legal sector, there's a growing category of companies building AI platforms to augment lawyers. These systems do more than help with generic writing tasks. They assist with contract review, diligence, research, matter management, and other workflows that are central to how legal work gets done.
Clinical trials, by contrast, are still largely stuck at the AI copilot stage.
Ask a sponsor or CRO where they use AI today in their clinical development process, and the answer is usually pretty basic. They've likely deployed generic AI copilots to help their team with general productivity tasks like cleaning up emails, summarizing long documents, and drafting trial materials. While incrementally useful, these products are not designed to deliver the step changes in trial velocity and cost structure promised by the mountain of AI hype in the industry.
There's More to a Clinical Trial Than Drafting an Email
The reality is that clinical trials don't usually fall behind because nobody can write a decent email. Rather, they do so because key operational workflows take too long to complete at the quality and scale trials demand, or because processes just haven't been designed well enough in the first place.
Let's consider what happens after a site monitoring visit. A CRA finds a series of findings and writes them up in the monitoring visit report. Writing this report can be supported by an AI copilot.
But then what?
The site will need to be informed of the finding, will need to respond, and the response will need to be reviewed. If it's incomplete, it needs a follow-up. That follow-up needs to be tracked until it's actually closed. If the same issue shows up at the next visit, someone needs to notice and track the trend. This complex, multi-step process extends beyond an AI copilot's ability to help speed up a CRA's content creation abilities. These are the types of workflows that overwhelm CRAs and slow down trials.
This same pattern of complex workflows slowing down trials shows up frequently in clinical operations:
In data management: Queries may sit unresponded to for 40 days. Someone eventually notices and acts on them, but by then the issues have either escalated or been added to by other sites in the study.
In TMF management: A missing delegation log is flagged. But who tracks that it was supposed to come in two months ago? Who follows up if it still hasn't arrived?
These aren't problems that can be solved with faster emails or a quickly summarized document. They're problems that require teams to address a workflow in its entirety, handling all of the steps, edge cases, follow-ups, and tracking required. AI copilots cannot solve this.
This is where agents come in.
Copilot vs. Agent
Many people know copilots and use them in their daily work. Examples include household names like ChatGPT from OpenAI, Copilot from Microsoft, and Claude from Anthropic. They are able to answer questions and have conversations like a human, usually with reference to documents that give them context on the task. For example, you ask it to draft an email, it writes some text, and you can then review it, edit it, and send it. Copilots help accelerate a person, helping them do their job faster, but they cannot automate complete workflows.
An agent is different. Agents are designed to execute multi-step workflows autonomously, only reporting to humans when a situation falls out of their scope of control. Agents use the same underlying technology as copilots, but layer on tools, evaluations, guardrails, and context databases that enable them to become experts at a specific workflow. By way of example, take the case of monitoring report review and finding resolution:
The copilot user experience:
- The team member asks the copilot to review the MVR for issues.
- Identified issues are reviewed by the team member to the best of their ability.
- If an action is required, they ask the copilot to draft an email, review it, and send it manually.
- They then need to manually track each of these findings, follow up on the emails, and ensure the findings are resolved.
If it takes 5 minutes to action each finding and 20 minutes of follow-up in the following days or weeks, across 30 reports with 10 findings each, that translates into 125 hours, or three weeks, of work.
The agent user experience:
- The team member logs into the Phases platform to see all 30 reports automatically reviewed, with all findings extracted with citations and actions suggested.
- They simply review the list of findings and suggested actions, approving, modifying, or denying each AI action suggestion. This takes 1-2 minutes per finding.
- Once approved, the AI agent drafts and sends emails, manages all follow-up, and reports back to the team member with the resolution outcomes autonomously.
If it takes 2 minutes to review and approve each AI-suggested action and 1 minute to review each resolution outcome, across 30 reports with 10 findings each, that translates into 15 hours of work: an 88% time saving compared to using a copilot.
Copilot vs. Agent workflow for reviewing monitoring reports and finding resolution, based on 30 reports with 10 findings each.
Why It's Hard to Build an Agent
Building an agent is much harder than building a copilot, especially in regulated environments. Unlike a copilot, an agent needs clearly defined guardrails to ensure it can only take actions it is explicitly permitted to take. Additionally, it needs tools to be defined to allow it to execute specific functions like interacting with external systems, performing calculations, or communicating with humans via email, text, or phone call. For instance, you could have a tool to QC a document in a TMF, a tool to flag a query that's been open too long, or a tool to email a site, and each tool needs clear inputs, outputs, boundaries, and logging.
Perhaps most importantly, an agent also requires a series of integrations to be useful. For instance, it can be integrated with your EDC, TMF, and CTMS to enhance its ability to operate autonomously and gather context to complete tasks. Without these integrations, an agent will lack the necessary context to perform the actions it requires effectively.
Now, you could use a generic agent here like Claude Cowork, but you'll quickly run into problems. For instance, you'll be lacking all sorts of integrations, the system itself will not be properly set up to log all actions in a manner that abides by 21 CFR Part 11 or EU Annex 11 compliance, and all the tools that would make an agent useful in a clinical operations setting simply wouldn't exist. You'd have to try to cobble it all together yourself, leading to an unvalidated and unstandardized system that lacks the guardrails and evaluations required to judge performance and output quality.
Conclusion
After speaking with 100+ clinical operations teams, it's clear that clinical research is lagging behind other industries in the adoption and deployment of AI that extends beyond copilots and the models themselves. Of course, this comes as no surprise given the regulated nature of the industry and the need to carefully validate all new technology for efficacy, compliance, and security.
If we really want to move the needle, however, there will need to be a conscious transition from copilots to agents, and specifically, purpose-built agents that integrate with the right systems, use the right tools, and abide by the correct regulations.
For now, agents are in the infancy of their adoption by CROs and sponsors. But for those who adopt them, they will have the potential to deliver step changes in trial efficiency and cost savings.
To learn more about agents and their impact on clinical research, reach out to the team at founders@phases.ai.