AI for research

Human in the loop, but which loop? Where oversight of AI actually helps in science

Where should a human actually review an AI’s scientific work, and what does the evidence say about that choice?

Not wherever the output looks important, which is where most checkpoints end up. A preregistered meta-analysis of 106 studies found human-AI combinations performing worse on average than the better of human or AI alone, and found the combination gained only where the human was already better at that specific judgement. That is a placement rule: put the checkpoint where the person has a real advantage, at actions that cannot be undone, and at boundaries where an error would go silent. Everything else is alert fatigue with extra steps.

By PharosBioPublished on 10 min read

Key takeaways

  • Human plus AI beat the better of either alone only when the human was already better at that judgement.
  • A checkpoint placed before the deciding evidence exists does not catch errors, it ratifies them.
  • Expert reviewers rated LLM research ideas more novel, then reversed once the ideas were executed.
  • Seventy teams analysing one dataset chose seventy different workflows and disagreed on which hypotheses held.
  • Clinicians override about 90% of drug interaction alerts; a checkpoint nobody reads is not oversight.
  • Automation bias affects experts as much as novices and cannot be trained away.
  • Place checkpoints at irreversibility, not at uncertainty, and at boundaries where errors go silent.
  • Set the constraints the system works inside rather than approving every move it proposes.

Start with the caveat, because it is also the point

The best evidence on this question does not come from biology. The controlled experiment on research ideas was run in natural language processing. The study of analytical variability was run on neuroimaging. The alert-fatigue data comes from hospital prescribing. Nobody has run the equivalent experiment in computational drug discovery.

These fields share the structure that matters: expert judgement applied to a technical artefact, at a point chosen by convention rather than by evidence. The findings transfer as hypotheses, not as proof, and the missing experiment is one worth running on our own field.

Definition

A human checkpoint is a point in an automated workflow where a person must review, approve or override before the work proceeds. Its value is not fixed. It depends entirely on whether, at that specific point, the person can see the information needed to make the judgement being asked of them.

Oversight is not free, and not always positive

The headline result is uncomfortable. Vaccaro, Almaatouq and Malone (2024) ran a preregistered systematic review and meta-analysis of 106 experimental studies reporting 370 effect sizes, and found that on average human-AI combinations performed significantly worse than the better of humans or AI alone. Losses concentrated in decision-making tasks; gains appeared in content-creation tasks.

The moderator analysis is what makes this actionable rather than merely dispiriting. When humans outperformed AI alone, the combination gained. When AI outperformed humans alone, the combination lost. Adding a person to a task they are worse at does not average the two. It drags the result down.

That converts a governance instinct into an empirical question. Before adding a checkpoint, the question is not “is this step important?” but “at this step, is the human better than the system at this judgement?” Those are different questions and they frequently have different answers.

The checkpoint that ratified instead of catching

The cleanest demonstration comes from a paired set of studies on scientific work itself.

Si, Yang and Hashimoto (2024) recruited over 100 NLP researchers to write research ideas and blind-review both human and LLM-generated ones. The LLM ideas were judged significantly more novel than the expert ideas, though slightly weaker on feasibility. Read at face value, that is a review checkpoint working: experts assessed the proposals and reached a verdict.

Then the same group executed them. Si, Hashimoto and Yang (2025)recruited 43 expert researchers, randomly assigned them ideas from either source, and had each spend over 100 hours actually implementing the idea and writing it up. The write-ups were then blind-reviewed again. The LLM ideas’ scores dropped significantly more than the human-written ones on all four metrics, novelty, excitement, effectiveness and overall, closing the ideation-stage gap. On many metrics the ranking flipped outright: human ideas now scored higher.

The reviewers were not weak. They were reviewing at a point where the deciding information did not yet exist.

This is the thesis in one experiment. The reviewers were qualified, motivated and blinded. What they lacked was evidence: at the ideation stage, empirical performance, baselines, ablations and resource requirements were speculative or invisible. A checkpoint placed there cannot discriminate on the criteria that ultimately decide the question, so it discriminates on what it can see, which is fluency and apparent originality. Language models are extremely good at those.

The generalisation is uncomfortable for a lot of deployed review processes. Any gate that asks a human to approve a plan before the plan has produced evidence is measuring the plan’s plausibility, not its correctness.

The reviewer is one draw from a wide distribution

A second objection to “have an expert sign off” is that expert sign-off is far less determinate than it sounds.

In the Neuroimaging Analysis Replication and Prediction Study,Botvinik-Nezer and colleagues (2020) gave 70 independent teams the same functional MRI dataset and the same nine pre-specified hypotheses. No two teams chose identical analysis workflows. That flexibility produced sizeable variation in which hypotheses came out significant, and it did so even between teams whose intermediate statistical maps were highly correlated. Two analysts can agree substantially about the data and still disagree about the answer.

The constructive half of that result matters just as much: meta-analytic approaches that aggregated across teams did yield significant consensus. Aggregation across analyses outperformed any single expert’s judgement.

For checkpoint design that argues for something other than more approval gates. If one expert’s analytical choices are a draw from a distribution that wide, the useful intervention is to run several defensible analyses and look at the spread, which is the multiverse approach, rather than to have one person initial the one path that happened to be taken. We made a related argument about aggregating agent runs in when agents disagree.

What a badly placed checkpoint decays into

Clinical decision support is the best-documented case of an oversight mechanism collapsing under its own volume, and the numbers are not marginal.

A 2024 systematic review and meta-analysis by Felisberto and colleagues found a pooled drug-drug interaction alert override rate of 90% (95% CI 85.6 to 95.0), drawn from 11 studies covering 570,776 prescriptions within a review of 16 articles. Asingle-hospital evaluation of a DDI system found 88.2% of 38,409 very severe interaction alerts overridden, with false positives traced to an overly broad screening interval and to the system not incorporating patient-specific characteristics.

The detail that makes this more than a cautionary tale: in that same evaluation, prescribers surveyed still found the system useful. The checkpoint was not resented. It was simply generating alerts faster than anyone could meaningfully evaluate them, and the rational response to a channel that is mostly false positives is to stop reading it.

The obvious remedy, better training, does not work. Parasuraman and Manzey (2010), reviewing the empirical literature on automation bias and complacency, reported that both appear in naive and expert participants alike, that complacency cannot be overcome with simple practice, and that automation bias cannot be prevented by training or instructions, producing both omission and commission errors when the aid is imperfect. This is a property of attention under multiple-task load, not a failure of diligence you can exhort people out of.

When oversight becomes assurance theatre

There is a sharper version of this critique, and it is worth confronting because it lands on the exact requirement most organisations are now trying to satisfy.

Green (2022) surveyed 41 policies prescribing human oversight of government algorithms and argued they suffer two flaws. First, the evidence suggests people are unable to perform the oversight functions the policies ask of them. Second, and consequently, the requirement legitimises deployment of faulty and controversial systems without addressing what is wrong with them, providing a false sense of security and letting vendors and agencies avoid accountability. His proposed remedy is a shift from human oversight to institutional oversight as the central regulatory mechanism.

This sits awkwardly beside regulation. The EU AI Act requires human oversight for high-risk systems under Article 14, and the NIST AI Risk Management Framework expects oversight to be demonstrable. Buyers therefore have to satisfy a requirement that a serious literature says is frequently performative.

The way out is not to argue against the requirement. It is to satisfy it with checkpoints that would survive an audit of whether they actually change decisions. A log showing that 90% of prompts were approved in under three seconds satisfies the letter of oversight and nothing else.

So where does the checkpoint go

Rules 1 and 4 follow the cited work. Rules 2 and 3 are ours, argued from the evidence rather than drawn from a single study.

Irreversibility, not uncertainty

The common design gates on model confidence: flag anything the system is unsure about. That produces exactly the DDI failure mode, because uncertainty is abundant and mostly cheap to revisit. Gate instead on what cannot be undone. A speculative intermediate result costs nothing to revisit; committing reagents, animal cohorts, or a trial arm does not. Those are different axes and only one of them tracks cost.

Boundaries where errors go silent

The other systematic error is checking only the final output. By then a wrong identifier mapping or an inappropriate threshold is invisible: it has been laundered into a conclusion that reads perfectly well. Errors have to be caught where they enter, which means checkpoints at chain boundaries, where one stage’s output becomes the next stage’s input. That is the argument we made at length about identifier resolution, and it generalises.

Set the box, do not approve every move

The self-driving laboratory field has already worked out useful vocabulary here, borrowing autonomy levels from autonomous vehicles. Levels run from 0, where humans do everything, to 5, full autonomy. At levels 2 and 3 the system runs design-build-test-learn cycles within constraints humans set, interpreting routine analyses and flagging anomalies for review. A 2024 review in Chemical Reviews reports that levels 2 and 3 account for the vast majority of self-driving laboratories demonstrated to date, and that a true level 5 remains unattained.

The useful part is the shape, not the label. Humans define the space the system may operate in and review what falls outside it. That is a checkpoint on the constraints, which a person can actually evaluate, rather than a checkpoint on every action, which they cannot.

Instead ofDo thisBecause
Approve each analysis stepApprove the analysis plan and the constraints, then review exceptionsStep approval is where automation bias and alert fatigue take hold
Flag every low-confidence outputFlag every irreversible action regardless of confidenceConfidence and cost are different axes; only one of them matters for damage
Review the final reportReview the identifier resolutions, thresholds and exclusions as they happenBy the report stage the error is invisible and reads as a finding
One expert signs offRun several defensible analyses and look at the spreadNARPS shows one expert's path is a draw from a wide distribution
Log that a human approvedLog what the human saw, changed and rejectedAn approval rate near 100% is evidence the checkpoint is not working

How we build this

Two design consequences follow, and they are the ones we act on.

The first is that validation should be a separate agent from the one doing the work, running continuously rather than as a terminal gate. Hydra runs a validator whose job is to check that the steps taken match the plan, that results are supported by what was actually retrieved, and that the analysis has not drifted from the question asked. That places checks at the chain boundaries rather than only on the output, which is rule three.

The second is that the human checkpoint we ask for is on the plan and the constraints, not on every step. A scientist sets the direction, the acceptance criteria and the limits; the system executes within them and surfaces what falls outside. That is rule four, and it is why the platform is built around planning and validation rather than a chat that asks for approval at each turn.

Both of those depend on provenance. A checkpoint is only real if the person can see what the system actually did: which databases and versions, which resolved queries, which records were excluded and why. Without that the reviewer is being asked to approve a summary, which is the ideation-stage problem wearing a different hat. Our work on predicting clinical outcome from in vivo data took the strict version of this: endpoints that cannot be resolved into a well-defined schema are excluded and flagged rather than coerced, which costs recall and buys a reviewable result.

The experiment nobody has run

Everything above transfers by analogy. The ideation-execution study was NLP. NARPS was neuroimaging. The override data is hospital prescribing. The direct experiment, taking a computational drug discovery workflow, varying where the human checkpoint sits, and measuring downstream decision quality, has not been published as far as we can find.

That is worth stating rather than papering over, and it is a reasonable thing for this field to go and do. Until then the honest position is that checkpoint placement should be treated as a design hypothesis with evidence behind it, not as a governance checkbox, and that the burden of proof sits with whoever wants to add a gate rather than with whoever questions one.

The uncomfortable summary: a checkpoint nobody reads is worse than no checkpoint, because it also supplies assurance. If your approval rate is near 100%, you do not have oversight. You have a log.

Hydra

Set the direction. Check the work that matters.

Hydra plans an analysis, runs it across roughly 100 preloaded scientific databases and 200+ codified skills, and validates every result with a separate agent, keeping the provenance you need to review it. Sign up and see what it shows you.

Glossary

TermWhat it means
Human checkpointA point where a person must review, approve or override before work proceeds
Automation biasOver-reliance on automated output, causing both missed errors and wrong actions
Automation complacencyReduced monitoring of an automated system under competing task load
Alert fatigueDesensitisation to warnings when most of them are false positives
Override rateThe share of alerts a user dismisses rather than acts on
Ideation-execution gapThe divergence between how an idea is judged before and after it is carried out
Analytical variabilityDifferent defensible workflows producing different conclusions from one dataset
Multiverse analysisRunning many defensible analyses and reporting the spread rather than one path
Autonomy levelsA 0 to 5 scale for how much of the design-execute-learn loop runs without a human
Institutional oversightGovernance at the organisational level rather than by an individual reviewer
IrreversibilityThe property that an action cannot be cheaply undone, which is what should gate it

Frequently asked questions

Does human oversight of AI actually improve results?

Only sometimes. A preregistered meta-analysis of 106 studies found human-AI combinations performed worse on average than the better of human or AI alone, with losses concentrated in decision-making. Combinations gained when the human already outperformed the AI at that task, and lost when the AI outperformed the human.

Where should a human checkpoint go in an AI research workflow?

At four places the evidence supports: where the human is measurably better at that specific judgement; at actions that are irreversible rather than merely uncertain; at boundaries where an error would become invisible downstream; and on the constraints the system operates within rather than on every individual step.

What is automation bias?

Automation bias is the tendency to over-rely on automated output, producing both errors of omission and commission when the aid is imperfect. Parasuraman and Manzey (2010) found it in novices and experts alike, and reported it cannot be prevented by training or instructions, which is why better reviewer training is not a fix.

Why do experts disagree when analysing the same scientific data?

Because analysis involves many defensible choices. In NARPS, 70 teams analysed one neuroimaging dataset against nine pre-specified hypotheses and no two chose identical workflows, producing sizeable variation in which hypotheses came out significant even between teams whose intermediate results were highly correlated.

Is required human oversight of AI just a formality?

It can become one. Green (2022) surveyed 41 policies mandating human oversight of government algorithms and argued they fail twice: people cannot reliably perform the oversight asked of them, and the requirement then legitimises the system without fixing it. He proposes institutional oversight instead.

Sources

  1. 01Vaccaro M, Almaatouq A, Malone T (2024). When combinations of humans and AI are useful: a systematic review and meta-analysis. Nature Human Behaviour8: 2293–2303. doi.org/10.1038/s41562-024-02024-1 Source for the 106 studies, 370 effect sizes, and the moderator result.
  2. 02Si C, Yang D, Hashimoto T (2024). Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers. arXiv:2409.04109. arxiv.org/abs/2409.04109
  3. 03Si C, Hashimoto T, Yang D (2025). The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas. arXiv:2506.20803. arxiv.org/abs/2506.20803 Source for the 43 executors, the 100+ hours, and the ranking flip. Note the author order differs from the 2024 paper.
  4. 04Botvinik-Nezer R, Holzmeister F, Camerer CF, et al. (2020). Variability in the analysis of a single neuroimaging dataset by many teams. Nature 582: 84–88. doi.org/10.1038/s41586-020-2314-9
  5. 05Felisberto M, Lima GS, Celuppi IC, et al. (2024). Override rate of drug-drug interaction alerts in clinical decision support systems: a brief systematic review and meta-analysis. Health Informatics Journal. doi.org/10.1177/14604582241263242 Pooled 90% (95% CI 85.6 to 95.0) from 11 studies and 570,776 prescriptions, within a review of 16 articles.
  6. 06Overall performance of a drug–drug interaction clinical decision support system: quantitative evaluation and end-user survey (2022). pmc.ncbi.nlm.nih.gov/articles/PMC8864797 Source for 88.2% of 38,409 very severe alerts overridden and the causes of false positives.
  7. 07Parasuraman R, Manzey DH (2010). Complacency and bias in human use of automation: an attentional integration. Human Factors 52(3): 381–410. doi.org/10.1177/0018720810376055
  8. 08Green B (2022). The flaws of policies requiring human oversight of government algorithms. Computer Law & Security Review 45: 105681. doi.org/10.1016/j.clsr.2022.105681 The count of 41 policies is taken from the author’s own abstract.
  9. 09Self-Driving Laboratories for Chemistry and Materials Science (2024). Chemical Reviews 124(16): 9633. doi.org/10.1021/acs.chemrev.4c00055 Source for the autonomy-level scale and for levels 2 and 3 covering the majority of demonstrated systems.
  10. 10Regulation (EU) 2024/1689 (AI Act), Article 14 on human oversight, and the NIST AI Risk Management Framework. Compliance dates for high-risk systems have been subject to amendment during 2026; verify the current position before planning to a deadline.

Related reading