Beyond Transcription Challenge
End-to-end audio models hallucinate nearly every clinical claim they generate. Can we fix that?
↗ Listen to a sample conversation from the dataset
About
Generating clinical notes directly from audio — skipping transcription — is faster, cheaper, and avoids cascading ASR errors. But today's end-to-end models hallucinate at alarming rates: on the Synth-DoPaCo dataset, 99–100% of their clinical claims are unsupported by the source audio, compared to just 21–23% for traditional transcribe-then-summarize pipelines. BeTraC is a shared evaluation challenge to close this gap — building end-to-end speech models that are actually faithful enough to trust in healthcare.
Tracks
Both tracks require open-weight models only and share the same constraint: no intermediate transcription.
Participation Rules
The full rules document includes approved model and dataset lists, detailed parameter counting rules (MoE, PLE/MatFormer, omni-model stripping), and submission requirements. Additional models or datasets may be proposed for inclusion by May 4, 2026 May 11, 2026+1 wk — contact betrac@googlegroups.com.
For questions or to register your team, contact betrac@googlegroups.com.
Dataset
Fully synthetic doctor-patient conversations generated with open-weight, permissively licensed models. Speaker identities are strictly disjoint across splits. Audio features two speakers, 66 ambient sound classes, room reverberation, and Opus compression artifacts.
🤗 View on Hugging FaceEvaluation Metrics
Post-Competition Analysis (Top 5 per Track)
Top 5 systems per track will undergo additional LLM-as-a-judge evaluation (Faithfulness, Coverage, Structure, Conciseness) and out-of-domain evaluation on real recorded OSCE interviews.
Leaderboard
| # | System | Pipeline | Concept F1 ↓ | C-Prec | C-Recall | ROUGE-2 | ROUGE-3 | ROUGE-L | Words |
|---|---|---|---|---|---|---|---|---|---|
| 1 | TalTech | end-to-end | 0.5429 | 0.5790 | 0.5177 | 0.3661 | 0.2497 | 0.4342 | 323 |
| 2 | NTT-HI-CS | end-to-end | 0.5397 | 0.5867 | 0.5069 | 0.3805 | 0.2646 | 0.4524 | 310 |
| 3 | KUSLP | end-to-end | 0.5152 | 0.5570 | 0.4882 | 0.3670 | 0.2529 | 0.4328 | 319 |
| 4 | JYVA | end-to-end | 0.5005 | 0.5430 | 0.4727 | 0.3513 | 0.2429 | 0.4120 | 323 |
| 5 | KIT-ISL-AI4LT | end-to-end | 0.4949 | 0.5459 | 0.4611 | 0.3601 | 0.2499 | 0.4323 | 291 |
| 6 | DirectSense | end-to-end | 0.4337 | 0.5410 | 0.3685 | 0.2487 | 0.1529 | 0.3315 | 241 |
| 7 | ASLP | end-to-end | 0.3965 | 0.3928 | 0.4134 | 0.2220 | 0.1342 | 0.3045 | 438 |
| 8 | IASP | end-to-end | 0.3616 | 0.4058 | 0.3354 | 0.1302 | 0.0799 | 0.1814 | 743 |
| 9 | Qwen2.5-Omni-3B ↗ code | end-to-end | 0.2449 | 0.2785 | 0.2257 | 0.0814 | 0.0290 | 0.1674 | 407 |
* Qwen2.5-Omni-3B is the official e2e baseline. Test-set results submitted by teams via the challenge submission process.
| # | System | Pipeline | Concept F1 ↓ | C-Prec | C-Recall | ROUGE-2 | ROUGE-3 | ROUGE-L | Words |
|---|---|---|---|---|---|---|---|---|---|
| 1 | TalTech | end-to-end | 0.5625 | 0.6015 | 0.5361 | 0.4009 | 0.2804 | 0.4680 | 312 |
| 2 | KUSLP | end-to-end | 0.5441 | 0.5956 | 0.5090 | 0.3927 | 0.2771 | 0.4586 | 311 |
| 3 | NTT-HI-CS | end-to-end | 0.5389 | 0.5794 | 0.5111 | 0.3771 | 0.2614 | 0.4482 | 315 |
| 4 | ASLP | end-to-end | 0.5058 | 0.5035 | 0.5182 | 0.3405 | 0.2267 | 0.4146 | 366 |
| 5 | Slate Lab | end-to-end | 0.4754 | 0.5102 | 0.4671 | 0.3000 | 0.2056 | 0.3566 | 471 |
| 6 | Qwen2.5-Omni-7B🧪 DEV ONLY ↗ code | end-to-end | 0.2572 | 0.3070 | 0.2302 | 0.0950 | 0.0343 | 0.1837 | 350 |
| 7 | Qwen2.5-Omni-3B ↗ code | end-to-end | 0.2449 | 0.2785 | 0.2257 | 0.0814 | 0.0290 | 0.1674 | 407 |
| 8 | Qwen3-Omni-30B-A3B-Instruct🧪 DEV ONLY ↗ code | end-to-end | 0.1879 | 0.1879 | 0.1964 | 0.0558 | 0.0175 | 0.1433 | 351 |
| 9 | Qwen3-Omni-30B-A3B-Thinking🧪 DEV ONLY ↗ code | end-to-end | 0.1645 | 0.1694 | 0.1675 | 0.0472 | 0.0123 | 0.1314 | 316 |
* Qwen2.5-Omni / Qwen3-Omni variants are the official e2e baselines. Test-set results submitted by teams via the challenge submission process.
🧪 DEV ONLY — evaluated on the dev set only, not yet run on the withheld test set.
| System | Pipeline | Concept F1 | C-Prec | C-Recall | ROUGE-2 | ROUGE-3 | ROUGE-L | Words |
|---|---|---|---|---|---|---|---|---|
| Whisper-large-v3 ASR + Qwen3-30B-A3B ↗ code | cascade | 0.2860 | 0.2964 | 0.2838 | 0.1174 | 0.0425 | 0.2243 | 261 |
| Qwen3-ASR-1.7B + Qwen3-30B-A3B ↗ code | cascade | 0.2772 | 0.2881 | 0.2741 | 0.1092 | 0.0374 | 0.2169 | 256 |
Schedule
Team