Research

OT Study Tests an AI Tool That Can Show Its Work

An OTJR study used 3,023 Japanese OT case reports to test traceable AI decision support, with clinical accuracy and patient outcomes still to be measured.

Artificial intelligenceGraphRAGClinical reasoningDecision supportOccupational therapy researchEvidenceAOTA ethics
Occupational therapist comparing handwritten assessment notes with a tablet in a rehabilitation clinic
Source-backed occupational therapy analysis from The OT Index News Desk.

Analysis based on the August 25 OTJR study, PubMed and Crossref publication records, Japanese Association of Occupational Therapists material on its case-report corpus, NIST work on biomedical source attribution, and AOTA's current AI policy.

Clinical AI has a credibility problem that begins after the answer appears. A fluent intervention plan may arrive in seconds; finding the facts that support each recommendation can take longer than writing the plan from scratch. A new OTJR study from researchers in Japan tested a narrower design: a system that retrieves occupational therapy cases through a knowledge graph before asking a language model to propose interventions. Across 3,023 case reports, that GraphRAG approach produced plans that stayed closer to the retrieved material than keyword, embedding, or random-case alternatives. The result is a useful engineering signal. It remains several steps away from a safe clinical product.

In brief

  • Researchers built a knowledge graph linking assessment findings, goals, and intervention plans across 3,023 reports in the Japanese Association of Occupational Therapists Case Report Corpus.
  • Using GPT-5-mini, the study compared five generation conditions. GraphRAG produced significantly higher RAGAS faithfulness scores than the other reference-based approaches.
  • The published evaluation measured whether generated claims stayed grounded in retrieved cases. It did not establish clinical accuracy, safety, therapist agreement, generalizability, or improved patient outcomes.

The case library comes first

The six-author team included occupational therapy researchers at Seijoh University and Nagoya University, along with an information-technology researcher affiliated with Seijoh and the Nagoya Institute of Technology's Artificial Intelligence Research Center. Their starting material was the Japanese Association of Occupational Therapists Case Report Corpus, an archive built from reports of actual OT practice.

A 2024 paper in the association's Asian Journal of Occupational Therapy says JAOT launched the registration system in 2005 to improve practice through case writing, document OT outcomes, and make the profession's work more visible. JAOT now says the original registration system has ended and the reports are available to members through its academic database. That history matters because the new system draws from a profession-specific record with a much narrower scope than general internet text.

The researchers organized the reports as a knowledge graph linking assessment findings, goals, and intervention plans. They placed 90% of the 3,023 cases in a reference set and held 10% back as a test set. Given assessment findings from a test case, GPT-5-mini generated a proposed intervention plan. The experiment asked whether a structured map of clinical relationships could help the model retrieve more useful precedents before it wrote.

Five ways to answer the same case

Each test case moved through five conditions. The model answered with no reference material, with randomly selected cases, with cases retrieved by keyword, with cases retrieved by embedding similarity, and with cases selected through GraphRAG. That final method uses the knowledge graph to follow relationships among clinical details. Shared words and mathematical similarity alone cannot capture all of those connections.

The generated plans were scored with the Retrieval-Augmented Generation Assessment faithfulness metric, commonly shortened to RAGAS faithfulness. The metric asks whether claims in an answer are supported by the context retrieved for that answer. Paired t-tests compared the models, and the OTJR abstract reports that GraphRAG scored significantly higher than the other approaches that used reference cases.

The abstract does not attach an effect size, score distribution, or error table to that conclusion. Even so, the comparison answers a concrete design question. A case library becomes more useful when its assessments, goals, and plans are connected in ways the retrieval system can follow. Dumping more records into a search box gives no such guarantee.

Faithfulness has a narrow job

Faithfulness is worth measuring because unsupported clinical prose can be hard to spot once it is polished. NIST's biomedical generative-retrieval work treats accurate source attribution as one way to reduce false statements from language models, especially when people are making clinical decisions or appraising research. A visible evidence trail gives the reviewer a place to begin.

Source support settles only one question. A recommendation can accurately reflect a retrieved case and still be poorly suited to the person in front of the therapist. The source case may carry different precautions, goals, culture, environment, service limits, or evidence quality. A case report records what one team did; it does not establish that the intervention caused the outcome or belongs in another plan of care.

The published abstract reports an automated faithfulness evaluation. It does not report occupational therapists rating the plans for appropriateness, checking contraindications, comparing the output with current guidelines, or measuring what happened to clients. Those tests belong between a promising retrieval result and clinical use.

The corpus sets the horizon

This system learned its clinical neighborhood from Japanese OT case reports. That focus is one of the study's strengths: the language of assessments, goals, and occupations comes from OT practice. It also defines the boundary of the result. Japanese service systems, documentation habits, scope rules, referral patterns, and cultural context shape the cases in the archive.

The reference and test cases came from the same corpus. An external test in another institution, practice setting, language, or country could expose different retrieval failures. U.S. teams would also need to account for payer requirements, state scope rules, privacy obligations, documentation standards, and source material that reflects the population they serve.

Case reports can supply precedents and questions. Clinical guidelines, systematic reviews, trials, current regulations, the occupational profile, and the therapist's own examination still carry separate jobs. A useful decision-support system should show which kind of source supports each claim and make its limits obvious at the moment of review.

What a clinic should demand before a pilot

Start with the trail. For every generated recommendation, the system should expose the retrieved case, the relevant excerpt, the date, the clinical setting, and the relationship the graph followed. Reviewers need to see when the tool found a close precedent, when it stretched a thin analogy, and when it found nothing. A citation that opens to an unrelated paragraph is a failure, even when the sentence sounds reasonable.

Then build a local challenge set before any live deployment. Include cases with contraindications, conflicting priorities, incomplete assessments, uncommon equipment, pediatric and adult distinctions, language differences, and goals that require environmental change. Have occupational therapists score clinical fit, missing precautions, invented facts, source quality, equity concerns, and time saved. Keep traceability and correctness as separate measures so a high score in one column cannot hide a failure in the other.

AOTA's AI policy leaves professional judgment with the practitioner and calls for accuracy, safety, transparency, evidence, equity, privacy, and accountability. Those requirements turn a technology demonstration into a governance project. Approved data handling, access controls, audit logs, version tracking, incident reporting, and a clear human owner for the final plan need to exist before identifiable client information reaches the system.

A trail worth following

The study's contribution is architectural. It organizes occupational therapy cases as connected clinical reasoning, then tests whether those connections help a language model stay grounded. The stronger faithfulness result says that design deserves a larger trial.

The next round should put clinicians into the evaluation, bring in outside cases, compare case reports with higher-level evidence, publish error categories, and measure whether the tool improves decisions without slowing or narrowing care. Patient and practitioner outcomes will matter more than the fluency of any generated plan.

Certainty is cheap on a screen. A visible trail gives the practitioner something concrete to inspect. This study has built that trail into the experiment. The profession now needs evidence about where it leads.

Decision use

How to use this analysis

Read the article first, then open the ranking table and related profiles to pressure-test the decision with source context.

Best OT Documentation Software1

Compare documentation systems with privacy, audit trails, workflow fit, and clinical review in view.

Open next step
Global Survey Finds AI Already in the OT Workday2

Read the companion survey on real-world OT use, disclosure, privacy, and governance.

Open next step
AJOT's New Platform Gives OT Readers a Short Window to Catch Up3

Use the remaining open-access window to review current occupational therapy evidence at its source.

Open next step