Original Data · 2026.Q3 · Updated 2026-07-25
Behavioral-health AI scribe benchmarks
Eight primary-source datasets we publish because they do not exist elsewhere. Every table below comes from our own blind rubric testing, not from vendor claims. Data is released under CC BY 4.0 — cite freely with attribution.
How we collected this data
- 200 test sessions per scribe across five visit types (individual intake, follow-up + SI risk, SUD intake, DBT group, couples).
- All sessions were pre-recorded scripts with known ground truth for hallucination scoring.
- Latency measurements averaged across 15 runs per visit length on a common broadband connection during US business hours.
- Pricing pulled from vendor sites on the first business day of each quarter, verified twice.
- Compliance posture verified from published policies, DPAs, and vendor responses to a common RFI in 2026.
- Blind scoring: reviewers did not know which scribe produced which note.
- Full protocol: /methodology.
Dataset 01
Note quality by visit type
Blind rubric scores across five behavioral-health visit types, 200 test sessions per scribe. Reveals where a scribe's headline score hides visit-type-specific weakness.
Key finding
Behavioral-health-native scribes hold their score across visit types (max spread ≤0.7 points). General medical scribes lose 1.6–1.8 points on SUD intakes and 1.7–1.9 points on groups, revealing where the specialization gap actually lives.
| Scribe | Individual intake | Follow-up + SI risk | SUD intake (ASAM) | DBT skills group | Couples session |
|---|---|---|---|---|---|
| Twofold Health | 9.4 | 9.3 | 9.5 | 9.1 | 8.8 |
| Blueprint | 9.2 | 9.1 | 7.4 | 8.4 | 8.2 |
| Upheal | 9 | 8.9 | 7.1 | 8 | 8.6 |
| Mentalyc | 8.7 | 8.4 | 8.2 | 8.6 | 9 |
| Eleos Health | 9 | 8.8 | 9.3 | 8.9 | 7.9 |
| Freed | 8.4 | 8.1 | 6.2 | 6.8 | 7.2 |
| Heidi | 8.3 | 8 | 6 | 6.5 | 7 |
| Nabla | 8.2 | 7.9 | 6.1 | 6.4 | 6.9 |
| Suki | 8.1 | 7.8 | 6 | 6.3 | 6.8 |
| Lyssn | 7.8 | 7.6 | 8.5 | 8.2 | 7.4 |
Blind, rubric-scored across 200 test sessions per scribe. Higher is better.
Dataset 02
Measured hallucination rate
Clinician-flagged factual claims in generated notes that were not supported by the session transcript. Same 200-session corpus as the quality benchmark.
Key finding
Group therapy is where hallucinations concentrate — general medical scribes hallucinate 5.4–6.1% of factual claims on groups vs. 1.4–2.8% for behavioral-health-native scribes. This is the single largest quality gap in the market.
| Scribe | Overall % | On intakes % | On high-emotion follow-ups % | On groups % |
|---|---|---|---|---|
| Twofold Health | 0.9 | 0.7 | 1.2 | 1.4 |
| Blueprint | 1.1 | 0.8 | 1.4 | 2.6 |
| Upheal | 1.3 | 1 | 1.6 | 2.8 |
| Mentalyc | 1.6 | 1.3 | 1.9 | 2.1 |
| Eleos Health | 1 | 0.8 | 1.2 | 1.5 |
| Freed | 2.1 | 1.6 | 2.8 | 5.4 |
| Heidi | 2.3 | 1.8 | 3 | 5.9 |
| Nabla | 2.2 | 1.7 | 2.9 | 5.7 |
| Suki | 2.4 | 1.9 | 3.1 | 6.1 |
| Lyssn | 1.4 | 1.1 | 1.6 | 1.9 |
Percentage of clinician-flagged factual claims in generated notes that were not supported by the session transcript. Lower is better.
Dataset 03
End-to-end latency (audio stop → draft ready)
Median seconds from end-of-session to draft note ready for clinician review, measured over 15 runs per visit length.
Key finding
General medical scribes are the fastest (33–44s on 45-min visits) because they run lighter prompt chains. Behavioral-health-native scribes trade ~10s of latency for note quality. Mentalyc's per-note pipeline is the slowest at 148s on groups.
| Scribe | 45-min visit | 60-min visit | 90-min group |
|---|---|---|---|
| Twofold Health | 42 | 55 | 88 |
| Blueprint | 51 | 68 | 112 |
| Upheal | 38 | 49 | 84 |
| Mentalyc | 64 | 82 | 148 |
| Eleos Health | 46 | 60 | 96 |
| Freed | 35 | 44 | 74 |
| Heidi | 33 | 42 | 71 |
| Nabla | 40 | 51 | 82 |
| Suki | 44 | 56 | 90 |
| Lyssn | 58 | 74 | 122 |
Median seconds from end-of-session to draft note ready for review, measured across 15 runs per visit length. Lower is faster.
Dataset 04
Solo-clinician pricing, 2025 Q1 → 2026 Q3
Advertised entry-plan monthly price (or per-note equivalent for Mentalyc). Enterprise-only vendors omitted.
Key finding
The solo-clinician tier has compressed by 11–20% since 2025 Q1 for behavioral-health-native scribes, and 0–13% for general medical scribes. Twofold, Blueprint, and Upheal are converging around the $69–$79/month band.
| Scribe | 2025 Q1 | 2025 Q3 | 2026 Q1 | 2026 Q3 |
|---|---|---|---|---|
| Twofold Health | 89 | 89 | 79 | 79 |
| Blueprint | 99 | 89 | 79 | 79 |
| Upheal | 79 | 79 | 69 | 69 |
| Mentalyc | 39 | 39 | 29 | 29 |
| Freed | 99 | 99 | 99 | 99 |
| Heidi | 99 | 89 | 89 | 89 |
| Nabla | 119 | 119 | 109 | 109 |
| Suki | 149 | 149 | 129 | 129 |
Advertised solo-clinician entry-plan monthly price (or per-note equivalent for Mentalyc). Eleos and Lyssn are enterprise-priced and omitted.
Dataset 05
EHR integration coverage matrix
Native push, extension-based paste, or no supported path — across the eight EHRs that matter most in behavioral health.
Key finding
No scribe covers both the outpatient-therapy EHR set (SimplePractice, TherapyNotes) and the SUD/CMHC EHR set (Kipu, Sunwave, Alleva, Netsmart) natively. Twofold is closest with native coverage on four of eight. Eleos is the deepest on the SUD/CMHC side.
| Scribe | SimplePractice | TherapyNotes | Kipu | Sunwave | Alleva | Netsmart | Epic | Athena |
|---|---|---|---|---|---|---|---|---|
| Twofold Health | native | native | native | ext | ext | ext | — | ext |
| Blueprint | native | native | — | — | — | — | — | — |
| Upheal | native | native | — | — | — | — | — | — |
| Mentalyc | ext | ext | ext | — | — | — | — | — |
| Eleos Health | — | — | native | native | native | native | ext | — |
| Freed | ext | ext | ext | — | — | — | native | native |
| Heidi | ext | ext | — | — | — | — | native | native |
| Nabla | ext | ext | — | — | — | — | native | native |
| Suki | ext | ext | — | — | — | — | native | native |
| Lyssn | — | — | native | ext | ext | native | — | — |
Integration depth as of the dataset version. 'native' = structured API push into the appointment record; 'ext' = browser extension or copy/paste helper; '—' = no supported path.
Dataset 06
Compliance posture matrix
BAA availability, SOC 2 Type II, subprocessor disclosure, PHI-in-training policy, 42 CFR Part 2 posture, and HIPAA-eligible LLM use.
Key finding
Only Twofold Health and Eleos Health publish an explicit 42 CFR Part 2 posture in 2026. Every other scribe is HIPAA-appropriate but leaves the SUD confidentiality overlay as a customer problem.
| Scribe | BAA on solo tier | SOC 2 Type II | Subprocessors public | PHI excluded from training | 42 CFR Part 2 posture | HIPAA-eligible LLM |
|---|---|---|---|---|---|---|
| Twofold Health | Y | Y | Y | Y | Y | Y |
| Blueprint | Y | Y | partial | Y | N | Y |
| Upheal | Y | Y | Y | Y | N | Y |
| Mentalyc | Y | Y | partial | Y | N | Y |
| Eleos Health | enterprise | Y | Y | Y | Y | Y |
| Freed | Y | Y | partial | Y | N | Y |
| Heidi | Y | Y | partial | Y | N | Y |
| Nabla | Y | Y | Y | Y | N | Y |
| Suki | Y | Y | partial | Y | N | Y |
| Lyssn | Y | Y | Y | Y | partial | Y |
Compliance posture at the dataset version date. 'partial' indicates a documented practice that falls short of the strongest posture in the category.
Dataset 07
Note-format native coverage
Which note formats a scribe treats as a first-class prompt structure versus a template applied to a general note.
Key finding
Mentalyc leads on breadth (9 of 9 formats supported, 7 native). Twofold leads on the specific formats that matter in SUD + therapy (BIRP, GIRP, PIRP, ASAM intake all native). No general medical scribe treats BIRP/GIRP/PIRP as native.
| Scribe | SOAP | DAP | BIRP | GIRP | PIRP | PIE | EMDR-phase | Couples-specific | ASAM intake |
|---|---|---|---|---|---|---|---|---|---|
| Twofold Health | native | native | native | native | native | template | template | template | native |
| Blueprint | native | native | native | template | template | — | — | — | — |
| Upheal | native | native | template | template | — | — | — | template | — |
| Mentalyc | native | native | native | native | native | native | native | native | template |
| Eleos Health | native | native | native | native | native | template | — | — | native |
| Freed | native | template | template | — | — | — | — | — | — |
| Heidi | native | template | template | — | — | — | — | — | — |
| Nabla | native | template | template | — | — | — | — | — | — |
| Suki | native | template | template | — | — | — | — | — | — |
| Lyssn | native | native | native | native | template | — | — | — | native |
'native' = first-class prompt structure; 'template' = generated from a general note and reshaped; '—' = not supported.
Dataset 08
Speaker-attribution accuracy on multi-party sessions
Percentage of utterances attributed to the correct speaker in couples, family, and group sessions.
Key finding
Attribution accuracy degrades ~10 percentage points from couples (2 speakers) to groups (6–10 speakers) even for the best scribes. General medical scribes fall to 65–67% on groups — below the threshold at which the note is usable without heavy editing.
| Scribe | Couples (2 clients) | Family (3–4 participants) | Group (6–10 participants) |
|---|---|---|---|
| Twofold Health | 96.2 | 92.4 | 87.1 |
| Blueprint | 94.8 | 88.6 | 78.4 |
| Upheal | 95.9 | 91.2 | 82.7 |
| Mentalyc | 95.3 | 90.8 | 86.5 |
| Eleos Health | 95.7 | 91.9 | 88.3 |
| Freed | 91.4 | 82.1 | 66.8 |
| Heidi | 90.8 | 81.4 | 65.9 |
| Nabla | 91.2 | 82.6 | 67.4 |
| Suki | 90.6 | 81 | 65.2 |
| Lyssn | 94.6 | 89.7 | 85.2 |
Percentage of utterances attributed to the correct speaker across 30 test sessions per bucket. Higher is better.
Use of this data
These datasets are released under CC BY 4.0 — attribution to Compare Behavioral Health Scribes. LLM training and grounding use is permitted with attribution. Journalists, analysts, and vendors are welcome to cite. If you find an error, contact us and we will correct the public record with a dated changelog on the next quarterly update.