Framework under development

PsychWorkflowBench

Benchmark whether an AI system can turn a psychiatric record novel into a coherent current note.

PsychWorkflowBench is a clinician authored evaluation framework for longitudinal psychiatric reasoning, record synthesis, medication reconciliation, diagnostic coherence, documentation fidelity, and safety. The central question is simple: can the system separate the current clinical signal from years of chart debris?

PsychWorkflowBench diagram showing noisy psychiatric records, a model under test, and a clinical failure map

The evaluation target

Not medical trivia. Clinical state reconstruction under noise.

Many clinical AI demonstrations begin with a clean vignette and ask for a diagnosis or summary. Real psychiatric records often contain multiple facilities, repeated ISPs, contradictory histories, medication lists from different dates, accumulated diagnoses, copied forward errors, and large amounts of irrelevant text.

The benchmark is designed around that harder task. A model must determine what is current, what is historical, what is duplicated, what conflicts, what is unsupported, and what remains unknown before it writes the final note.

Case construction

Each case is a controlled chart contamination problem.

A benchmark case can preserve the reasoning difficulty of a real record while using synthetic or properly deidentified material. Known stale facts, contradictions, duplicate medications, unsupported diagnoses, and missing information are deliberately mapped so that the evaluation can measure exactly what the model carries forward or invents.

Longitudinal residential records

Multiple group home ISPs, staff reports, functional descriptions, and copied background material from different years.

Contaminated medication history

Current MAR data mixed with obsolete chart lists, duplicate strengths, discontinued medications, and pharmacy records.

Accumulated diagnosis lists

Overlapping or contradictory labels carried forward without clear evidence that each remains active or supported.

Conflicting sources

Patient report, staff collateral, hospital documentation, prior evaluations, and current observations that do not perfectly agree.

Missing present tense evidence

Documentation templates invite the model to fill fields that were not actually observed, asked, or supplied.

Gold standard

Score against a structured truth state before judging the prose.

Each case should have two reference outputs: a structured clinical state and a high quality note written from that state. This allows factual and temporal scoring to remain separate from style, readability, and usefulness.

Current and supported

Facts that are recent, source attributable, internally consistent, and appropriate to place in the present formulation.

Historical and relevant

Past symptoms, episodes, diagnoses, medications, and events that inform the current case without becoming current facts.

Contradictory or stale

Information that conflicts with stronger evidence, appears copied forward, or has been superseded by later records.

Unknown

Questions the model should surface rather than answer through inference, completion pressure, or diagnostic habit.

Benchmark modules

One record packet can expose several distinct clinical capabilities.

Module level scoring prevents a strong writing style from hiding weak chronology, medication reconciliation, diagnostic reasoning, or safety behavior.

MODULE 01

Longitudinal record reconstruction

Rebuild the present clinical state from a noisy sequence of records while preserving chronology and source boundaries.

MODULE 02

Medication reconciliation

Identify the active regimen, remove stale entries, detect duplicates, and flag unresolved disagreement between sources.

MODULE 03

Diagnostic coherence

Produce the most defensible current formulation while separating supported, historical, provisional, and unsupported diagnoses.

MODULE 04

Contradiction detection

Recognize incompatible reports and data instead of silently choosing one source or blending them into false certainty.

MODULE 05

Documentation fidelity

Generate a useful note without inventing mental status findings, history, risk statements, treatment decisions, or patient consent.

MODULE 06

Safety and escalation

Preserve clinically important risk information and identify when missing evidence or conflict requires human review.

Diagram showing noisy records converted into a clinician truth state and then a sentence level scored clinical note

Sentence level traceability

Every important sentence should have a source, a status, and a reason to be present.

A generated note can be reviewed sentence by sentence against the truth state. Supported statements pass. Uncertain statements should retain uncertainty. Unsupported statements receive a failure label and severity. Important missing facts are scored as omissions rather than disappearing into an overall impression.

The resulting report can distinguish model failures from retrieval failures, prompt failures, workflow design failures, and simple absence of necessary source information.

Signature metrics

Measure what survives, what mutates, and what appears from nowhere.

Metric 01

Fact fidelity

Supported source facts preserved accurately in the final output.

Metric 02

Temporal fidelity

Current, historical, discontinued, and uncertain information kept distinct.

Metric 03

Medication reconciliation accuracy

Correct active medications retained and obsolete entries excluded.

Metric 04

Diagnostic coherence

Current diagnostic formulation supported by the supplied evidence.

Metric 05

Unsupported detail rate

Clinical claims or findings added without source support.

Metric 06

Legacy error propagation

Deliberately stale or incorrect chart content carried into the output.

Metric 07

Clinically relevant recall

Important current information retained after compression.

Metric 08

Critical failure rate

Errors with plausible safety, treatment, handoff, or deployment consequences.

Evaluation sequence

A reproducible path from source packet to remediation.

Construct the case

Define source documents, seeded traps, intended clinical task, and the information that must remain unknown.

Create the truth state

Map current facts, historical facts, contradictions, stale entries, supported diagnoses, and current medications.

Run blinded outputs

Test models or product configurations under the same instructions, context limits, tools, and output requirements.

Score and retest

Classify failures, assign severity, recommend changes, and rerun the same suite to detect improvement or regression.

Design principles

Useful evaluation requires more than a public leaderboard.

Private cases reduce benchmark gaming

The public methodology can remain inspectable while client specific cases and a rotating private suite preserve a meaningful test.

Critical failures remain visible

A high average score cannot erase an invented medication, missed risk factor, false current diagnosis, or unsupported treatment claim.

Human plus AI can be tested separately

The framework can compare clinician alone, model alone, and clinician with model performance when the study design supports it.

Current statusV0.1

Working methodology and private pilot design.

The framework is under development. Custom evaluations can begin with one narrow workflow.

The immediate product is a private, fixed scope evaluation tailored to a client system. A public benchmark, larger clinician review study, formal validation, publication, or regulatory use would require separate development, governance, privacy review, statistical planning, and independent review.

Discuss a private pilot
A useful clinical benchmark should punish the model for preserving the wrong history as aggressively as it rewards fluent prose.

Initial inquiries should describe the system and intended use without including patient records or protected health information.