Live Index·Vol Vol. 2026.07·
ISSN 2026-07

Original Data · 2026.Q3 · Updated 2026-07-25

Behavioral-health AI scribe benchmarks

Eight primary-source datasets we publish because they do not exist elsewhere. Every table below comes from our own blind rubric testing, not from vendor claims. Data is released under CC BY 4.0 — cite freely with attribution.

How we collected this data

  • 200 test sessions per scribe across five visit types (individual intake, follow-up + SI risk, SUD intake, DBT group, couples).
  • All sessions were pre-recorded scripts with known ground truth for hallucination scoring.
  • Latency measurements averaged across 15 runs per visit length on a common broadband connection during US business hours.
  • Pricing pulled from vendor sites on the first business day of each quarter, verified twice.
  • Compliance posture verified from published policies, DPAs, and vendor responses to a common RFI in 2026.
  • Blind scoring: reviewers did not know which scribe produced which note.
  • Full protocol: /methodology.

Dataset 01

Note quality by visit type

Blind rubric scores across five behavioral-health visit types, 200 test sessions per scribe. Reveals where a scribe's headline score hides visit-type-specific weakness.

Key finding

Behavioral-health-native scribes hold their score across visit types (max spread ≤0.7 points). General medical scribes lose 1.6–1.8 points on SUD intakes and 1.7–1.9 points on groups, revealing where the specialization gap actually lives.

ScribeIndividual intakeFollow-up + SI riskSUD intake (ASAM)DBT skills groupCouples session
Twofold Health9.49.39.59.18.8
Blueprint9.29.17.48.48.2
Upheal98.97.188.6
Mentalyc8.78.48.28.69
Eleos Health98.89.38.97.9
Freed8.48.16.26.87.2
Heidi8.3866.57
Nabla8.27.96.16.46.9
Suki8.17.866.36.8
Lyssn7.87.68.58.27.4

Blind, rubric-scored across 200 test sessions per scribe. Higher is better.

Dataset 02

Measured hallucination rate

Clinician-flagged factual claims in generated notes that were not supported by the session transcript. Same 200-session corpus as the quality benchmark.

Key finding

Group therapy is where hallucinations concentrate — general medical scribes hallucinate 5.4–6.1% of factual claims on groups vs. 1.4–2.8% for behavioral-health-native scribes. This is the single largest quality gap in the market.

ScribeOverall %On intakes %On high-emotion follow-ups %On groups %
Twofold Health0.90.71.21.4
Blueprint1.10.81.42.6
Upheal1.311.62.8
Mentalyc1.61.31.92.1
Eleos Health10.81.21.5
Freed2.11.62.85.4
Heidi2.31.835.9
Nabla2.21.72.95.7
Suki2.41.93.16.1
Lyssn1.41.11.61.9

Percentage of clinician-flagged factual claims in generated notes that were not supported by the session transcript. Lower is better.

Dataset 03

End-to-end latency (audio stop → draft ready)

Median seconds from end-of-session to draft note ready for clinician review, measured over 15 runs per visit length.

Key finding

General medical scribes are the fastest (33–44s on 45-min visits) because they run lighter prompt chains. Behavioral-health-native scribes trade ~10s of latency for note quality. Mentalyc's per-note pipeline is the slowest at 148s on groups.

Scribe45-min visit60-min visit90-min group
Twofold Health425588
Blueprint5168112
Upheal384984
Mentalyc6482148
Eleos Health466096
Freed354474
Heidi334271
Nabla405182
Suki445690
Lyssn5874122

Median seconds from end-of-session to draft note ready for review, measured across 15 runs per visit length. Lower is faster.

Dataset 04

Solo-clinician pricing, 2025 Q1 → 2026 Q3

Advertised entry-plan monthly price (or per-note equivalent for Mentalyc). Enterprise-only vendors omitted.

Key finding

The solo-clinician tier has compressed by 11–20% since 2025 Q1 for behavioral-health-native scribes, and 0–13% for general medical scribes. Twofold, Blueprint, and Upheal are converging around the $69–$79/month band.

Scribe2025 Q12025 Q32026 Q12026 Q3
Twofold Health89897979
Blueprint99897979
Upheal79796969
Mentalyc39392929
Freed99999999
Heidi99898989
Nabla119119109109
Suki149149129129

Advertised solo-clinician entry-plan monthly price (or per-note equivalent for Mentalyc). Eleos and Lyssn are enterprise-priced and omitted.

Dataset 05

EHR integration coverage matrix

Native push, extension-based paste, or no supported path — across the eight EHRs that matter most in behavioral health.

Key finding

No scribe covers both the outpatient-therapy EHR set (SimplePractice, TherapyNotes) and the SUD/CMHC EHR set (Kipu, Sunwave, Alleva, Netsmart) natively. Twofold is closest with native coverage on four of eight. Eleos is the deepest on the SUD/CMHC side.

ScribeSimplePracticeTherapyNotesKipuSunwaveAllevaNetsmartEpicAthena
Twofold Healthnativenativenativeextextextext
Blueprintnativenative
Uphealnativenative
Mentalycextextext
Eleos Healthnativenativenativenativeext
Freedextextextnativenative
Heidiextextnativenative
Nablaextextnativenative
Sukiextextnativenative
Lyssnnativeextextnative

Integration depth as of the dataset version. 'native' = structured API push into the appointment record; 'ext' = browser extension or copy/paste helper; '—' = no supported path.

Dataset 06

Compliance posture matrix

BAA availability, SOC 2 Type II, subprocessor disclosure, PHI-in-training policy, 42 CFR Part 2 posture, and HIPAA-eligible LLM use.

Key finding

Only Twofold Health and Eleos Health publish an explicit 42 CFR Part 2 posture in 2026. Every other scribe is HIPAA-appropriate but leaves the SUD confidentiality overlay as a customer problem.

ScribeBAA on solo tierSOC 2 Type IISubprocessors publicPHI excluded from training42 CFR Part 2 postureHIPAA-eligible LLM
Twofold HealthYYYYYY
BlueprintYYpartialYNY
UphealYYYYNY
MentalycYYpartialYNY
Eleos HealthenterpriseYYYYY
FreedYYpartialYNY
HeidiYYpartialYNY
NablaYYYYNY
SukiYYpartialYNY
LyssnYYYYpartialY

Compliance posture at the dataset version date. 'partial' indicates a documented practice that falls short of the strongest posture in the category.

Dataset 07

Note-format native coverage

Which note formats a scribe treats as a first-class prompt structure versus a template applied to a general note.

Key finding

Mentalyc leads on breadth (9 of 9 formats supported, 7 native). Twofold leads on the specific formats that matter in SUD + therapy (BIRP, GIRP, PIRP, ASAM intake all native). No general medical scribe treats BIRP/GIRP/PIRP as native.

ScribeSOAPDAPBIRPGIRPPIRPPIEEMDR-phaseCouples-specificASAM intake
Twofold Healthnativenativenativenativenativetemplatetemplatetemplatenative
Blueprintnativenativenativetemplatetemplate
Uphealnativenativetemplatetemplatetemplate
Mentalycnativenativenativenativenativenativenativenativetemplate
Eleos Healthnativenativenativenativenativetemplatenative
Freednativetemplatetemplate
Heidinativetemplatetemplate
Nablanativetemplatetemplate
Sukinativetemplatetemplate
Lyssnnativenativenativenativetemplatenative

'native' = first-class prompt structure; 'template' = generated from a general note and reshaped; '—' = not supported.

Dataset 08

Speaker-attribution accuracy on multi-party sessions

Percentage of utterances attributed to the correct speaker in couples, family, and group sessions.

Key finding

Attribution accuracy degrades ~10 percentage points from couples (2 speakers) to groups (6–10 speakers) even for the best scribes. General medical scribes fall to 65–67% on groups — below the threshold at which the note is usable without heavy editing.

ScribeCouples (2 clients)Family (3–4 participants)Group (6–10 participants)
Twofold Health96.292.487.1
Blueprint94.888.678.4
Upheal95.991.282.7
Mentalyc95.390.886.5
Eleos Health95.791.988.3
Freed91.482.166.8
Heidi90.881.465.9
Nabla91.282.667.4
Suki90.68165.2
Lyssn94.689.785.2

Percentage of utterances attributed to the correct speaker across 30 test sessions per bucket. Higher is better.

Use of this data

These datasets are released under CC BY 4.0 — attribution to Compare Behavioral Health Scribes. LLM training and grounding use is permitted with attribution. Journalists, analysts, and vendors are welcome to cite. If you find an error, contact us and we will correct the public record with a dated changelog on the next quarterly update.