Beyond Transcription Challenge
End-to-end audio models hallucinate nearly every clinical claim they generate. Can we fix that?
↗ Listen to a sample conversation from the dataset
About
Generating clinical notes directly from audio — skipping transcription — is faster, cheaper, and avoids cascading ASR errors. But today's end-to-end models hallucinate at alarming rates: on the Synth-DoPaCo dataset, 99–100% of their clinical claims are unsupported by the source audio, compared to just 21–23% for traditional transcribe-then-summarize pipelines. BeTraC is a shared evaluation challenge to close this gap — building end-to-end speech models that are actually faithful enough to trust in healthcare.
Tracks
Both tracks require open-weight models only and share the same constraint: no intermediate transcription.
Participation Rules
The full rules document includes approved model and dataset lists, detailed parameter counting rules (MoE, PLE/MatFormer, omni-model stripping), and submission requirements. Additional models or datasets may be proposed for inclusion by May 4, 2026 May 11, 2026+1 wk — contact betrac@googlegroups.com.
For questions or to register your team, contact betrac@googlegroups.com.
Dataset
Fully synthetic doctor-patient conversations generated with open-weight, permissively licensed models. Speaker identities are strictly disjoint across splits. Audio features two speakers, 66 ambient sound classes, room reverberation, and Opus compression artifacts.
🤗 View on Hugging FaceEvaluation Metrics
Post-Competition Analysis (Top 5 per Track)
Top 5 systems per track will undergo additional LLM-as-a-judge evaluation (Faithfulness, Coverage, Structure, Conciseness) and out-of-domain evaluation on real recorded OSCE interviews.
Leaderboard
| # | System | Pipeline | Concept F1 ↓ | C-Prec | C-Recall | ROUGE-2 | ROUGE-3 | ROUGE-L | Words |
|---|---|---|---|---|---|---|---|---|---|
| 1 | TalTech | end-to-end | 0.5429 | 0.5790 | 0.5177 | 0.3661 | 0.2497 | 0.4342 | 323 |
| 2 | NTT-HI-CS | end-to-end | 0.5397 | 0.5867 | 0.5069 | 0.3805 | 0.2646 | 0.4524 | 310 |
| 3 | KUSLP | end-to-end | 0.5152 | 0.5570 | 0.4882 | 0.3670 | 0.2529 | 0.4328 | 319 |
| 4 | JYVA | end-to-end | 0.5005 | 0.5430 | 0.4727 | 0.3513 | 0.2429 | 0.4120 | 323 |
| 5 | KIT-ISL-AI4LT | end-to-end | 0.4949 | 0.5459 | 0.4611 | 0.3601 | 0.2499 | 0.4323 | 291 |
| 6 | DirectSense | end-to-end | 0.4337 | 0.5410 | 0.3685 | 0.2487 | 0.1529 | 0.3315 | 241 |
| 7 | ASLP | end-to-end | 0.3965 | 0.3928 | 0.4134 | 0.2220 | 0.1342 | 0.3045 | 438 |
| 8 | IASP | end-to-end | 0.3616 | 0.4058 | 0.3354 | 0.1302 | 0.0799 | 0.1814 | 743 |
| 9 | Qwen2.5-Omni-3B ↗ code | end-to-end | 0.3089 | 0.343 | 0.295 | 0.116 | 0.050 | 0.189 | 817.4 |
* Qwen2.5-Omni-3B is the official e2e baseline. Test-set results submitted by teams via the challenge submission process.
| # | System | Pipeline | Concept F1 ↓ | C-Prec | C-Recall | ROUGE-2 | ROUGE-3 | ROUGE-L | Words |
|---|---|---|---|---|---|---|---|---|---|
| 1 | TalTech | end-to-end | 0.5625 | 0.6015 | 0.5361 | 0.4009 | 0.2804 | 0.4680 | 312 |
| 2 | KUSLP | end-to-end | 0.5441 | 0.5956 | 0.5090 | 0.3927 | 0.2771 | 0.4586 | 311 |
| 3 | NTT-HI-CS | end-to-end | 0.5389 | 0.5794 | 0.5111 | 0.3771 | 0.2614 | 0.4482 | 315 |
| 4 | ASLP | end-to-end | 0.5058 | 0.5035 | 0.5182 | 0.3405 | 0.2267 | 0.4146 | 366 |
| 5 | Slate Lab | end-to-end | 0.4754 | 0.5102 | 0.4671 | 0.3000 | 0.2056 | 0.3566 | 471 |
| 6 | Qwen2.5-Omni-7B ↗ code | end-to-end | 0.3261 | 0.379 | 0.296 | 0.140 | 0.061 | 0.224 | 448.2 |
| 7 | Qwen3-Omni-30B-A3B-Instruct ↗ code | end-to-end | 0.3178 | 0.267 | 0.403 | 0.146 | 0.064 | 0.231 | 563.1 |
| 8 | Qwen2.5-Omni-3B ↗ code | end-to-end | 0.3089 | 0.343 | 0.295 | 0.116 | 0.050 | 0.189 | 817.4 |
| 9 | Qwen3-Omni-30B-A3B-Thinking ↗ code | end-to-end | 0.2854 | 0.239 | 0.363 | 0.133 | 0.055 | 0.219 | 509.8 |
* Qwen2.5-Omni / Qwen3-Omni variants are the official e2e baselines. Test-set results submitted by teams via the challenge submission process.
| System | Pipeline | Concept F1 | C-Prec | C-Recall | ROUGE-2 | ROUGE-3 | ROUGE-L | Words |
|---|---|---|---|---|---|---|---|---|
| Whisper-large-v3 ASR + Qwen3-30B-A3B ↗ code | cascade | 0.2879 | 0.2964 | 0.2838 | 0.1174 | 0.0425 | 0.2243 | 261 |
| Qwen3-ASR-1.7B + Qwen3-30B-A3B ↗ code | cascade | 0.2811 | 0.2881 | 0.2741 | 0.1092 | 0.0374 | 0.2169 | 256 |
Schedule
Team