Research & validation

Published evidence for AIPRA as systematic review software—including a non-inferiority study, a narrative comparison with Covidence and Rayyan, and peer-reviewed outputs. Limitations are stated plainly.

Evidence-synthesis organizations increasingly expect authors to justify AI tool use with citations to performance evaluations—not marketing claims alone. The November 2025 joint position statement from Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence requires transparent reporting of AI systems, including evidence that a tool is methodologically sound for the intended task, with human oversight. The RAISE framework on OSF operationalizes those expectations for developers and users.

This page collects AIPRA-related publications you can cite when documenting tool choice. For manuscript-ready wording, use the AI disclosure methods paragraph.

Non-inferiority study (medRxiv)

AIPRA vs human-led systematic review pipeline (same question)

Two independent pipelines (human-led vs AIPRA) each produced a full systematic review manuscript on: “What is the role of large language models in glaucoma diagnosis?” Blinded domain experts rated overall quality. Mean total scores: human 74.7%, AIPRA 65.3%. Mean difference (AIPRA − Human) −9.3% (95% CI −18.8% to 0.0%), meeting the pre-specified non-inferiority criterion. AIPRA completed the workflow in approximately 2 hours versus about 1 month for the human pipeline.

Limitations (from the study itself). Domain means were identical for query development; the human-led pipeline scored higher in screening, field selection, full-text extraction, and manuscript writing. Appropriate human oversight remains essential—especially for screening and extraction. Do not treat time savings as a substitute for methodological judgment.

Read the medRxiv preprint

Platform comparison (High Yield Medical Reviews)

Systematic Review Management Platforms Powered by Artificial Intelligence: A Narrative Review

Narrative overview comparing Rayyan, Covidence, and AIPRA. Rayyan is positioned as a widely used screening-oriented platform; Covidence as a structured institutional workbench for screening through quality assessment; AIPRA as an AI-augmented platform spanning question formulation through manuscript drafting. The review emphasizes that platforms support—but do not replace—human oversight in evidence synthesis.

Read the HYMR article

Peer-reviewed output

Neurologic Clinics publication

Peer-reviewed publication supporting evidence synthesis conducted with the platform workflow.

View abstract

How we position AI claims

  • We prioritize recall/sensitivity and auditability over “review in minutes” messaging.
  • We do not claim automated risk-of-bias assessment as a solved problem; published LLM accuracy on selective-reporting judgments remains limited.
  • Authors remain responsible for protocol compliance, inclusion decisions, and manuscript accuracy—consistent with the 2025 multi-organization AI position statement.

Compare tools on /compare or read the systematic review software guide.