About

A professional profile with fewer adjectives and more receipts.

Stephen Thomas McCarthy is a physician assistant and clinical AI red team consultant based in Allentown, Pennsylvania. His background includes inpatient and outpatient psychiatric care, addiction treatment, medication management, telehealth, clinical leadership, documentation review, and work with diverse patient populations.

Stephen McCarthy

Background

McCarthy earned a Master of Science in Physician Assistant Studies and a Bachelor of Science in Medical Studies from DeSales University. His clinical work has centered on the practical problems of psychiatric assessment, medication management, impairment, documentation quality, longitudinal record review, and care coordination.

His consulting work applies frontline psychiatric judgment to clinical AI systems. The current focus is red team testing and benchmarking of longitudinal record synthesis, medication reconciliation, diagnostic coherence, contradiction handling, documentation fidelity, unsupported assertions, and safety behavior.

PsychWorkflowBench is the working name for an evaluation framework under development. It is designed to test whether a system can turn noisy psychiatric records into a coherent current representation without copying obsolete information or inventing details.

His writing focuses on psychiatric language: what diagnoses measure, what they help clinicians do, and what they cannot explain without independent causal evidence. The central concern is precision, not denial. Symptoms, distress, and impairment can be real even when the ontology of a diagnostic category remains unsettled.

Stephen McCarthy is licensed in Pennsylvania and Utah.

Professional method

Separate the layers

Source state

Current, historical, contradictory, unknown

Reviews begin by separating what is current and supported from what is old, copied, conflicting, or unresolved.

Clinical review

Observation before conclusion

Findings distinguish what the submitted material shows, what can reasonably be inferred, what remains uncertain, and what the system must not invent.

Communication

Arguments labeled as arguments

Reports and essays distinguish a thesis from a verified fact, cite their sources, and make uncertainty visible rather than decorative.

Clinical AI evaluation

Red team testing and benchmarking

Private evaluations can test one defined behavioral health workflow, a set of deidentified model outputs, or a model configuration against clinician authored cases and scoring criteria. The result is a severity ranked failure map, source linked findings, remediation priorities, and objective retest criteria.

Current writing focus

Psychiatric ontology without the fog machine

The newest essay argues that a syndrome label may be useful while remaining different from an independently observed disease process. It uses AuDHD to show why co occurrence, classification, and causation must not be collapsed into a single word.

Read the essay