Tessact is a reference, not ground truth. Every score here measures agreement with Tessact’s 20 Sep output. Where a model and Tessact disagree, either may be right. Human gold labels (runbook Phase 2) are needed before anyone can say a system is better than Tessact.
The title
- Title
- Meyebela, Season 1, Episode 1: “Roommates forever?” (Bengali drama; Kolkata flat-share, colloquial Bengali with English code-mix). Watch on hoichoi.tv/shows/meyebela.
- Media
- 762.0 s (12 min 42 s) · H.264 1920×1080 25 fps, AAC · 1.87 GB master
- Fingerprint
SHA-256 fe695db1…3603; the same bytes Tessact processed on 20 Sep and again in the controlled reingest on 27/28 Sep- Speech
- VAD found 87 speech regions, merged into 43 chunks totalling 300.5 s. Every open model transcribed exactly these chunks, so only the model differs.
- Hardware
- Apple M5 Pro, 64 GB. Speeds are laptop speeds; the AWS GPU run is still to come.
Scores
Three views, because whole-file comparison is dominated by conventions rather than recognition (see why). CER and WER are against Tessact 20 Sep; lower means closer to Tessact.
| Configuration | Group | × real time | CER · whole file | CER · chunk, Bengali script | CER · pure-Bengali chunks | WER · pure-Bengali | Licence |
|---|---|---|---|---|---|---|---|
| Loading scores… | |||||||
Pure-Bengali view: the 8 speech chunks whose Tessact text has no Latin letters (680 characters). Small sample; indicative only.
Diff viewer
Each row is one speech chunk, or a caption some system placed outside detected speech (italic). Pick a base column; every other column highlights words it adds relative to the base (added) and, optionally, base words it misses (missed). Collapse a column with −; open details in a header for run metadata. The URL keeps your selection, so a link opens the same view.
Why whole-file scores mislead
- Non-speech captions
- 14 of Tessact’s 92 cues (444 characters) sit mostly outside detected speech, including a 65 s cue over the music intro and a 139 s “Good morning.” cue. Filter captions outside speech to see them.
- Code-mix script
- 23% of Tessact’s characters are Latin: English words written in English. Every open model writes them phonetically in Bengali script (“Time bound” → “টাইম বাউন্ড”). That counts as an error without being a mishearing. Tick Bengali script only to hide Latin text.
- Uncaptioned English
- Some chunks contain English speech Tessact did not caption at all while the models did (e.g. a 10 s line where Tessact has only “মানে মানে”). These count against the models but may be Tessact omissions.
- Spelling
- Colloquial variants such as পারবোনি / পারবো নি, আনব / আনবো, দাও / দেও differ in characters, not meaning.
What the diff shows
- Lead model
- BengaliAI regional Whisper-medium (Apache-2.0) is closest to Tessact on pure Bengali (CER 0.17 with beam 5, 0.19 greedy) and never loops. On clean dialogue it is often near-identical; one 12 s chunk differs by 2%, in spelling only.
- Fine-tunes win
- Bengali fine-tuned medium-size models beat multilingual Whisper large-v3 by a wide margin at a fraction of the compute.
- Fast second opinion
- IndicConformer 600M runs ~120–140× real time on CPU with mid-table accuracy, which makes it useful for ensembling or a cheap first pass.
- Rejected
- Whisper large-v3-turbo loops on this audio (21–23 looping segments). Whisper large-v3 without VAD hallucinates repeated words over music.
- Real errors
- Place names (উলুবড়িয়া → কুলবাড়িয়া) and very short, noisy chunks.
- Vendor noise floor
- Tessact vs its own reingest of identical bytes: CER 0.04–0.06, scene count 6 → 8. Any system scoring near that is within the vendor’s own run-to-run variation.
Path to match or beat Tessact
- Fix the code-mix convention (English in Latin vs Bengali script) and add a transliteration pass.
- Spelling-tolerant scoring, so recognition is measured rather than orthography.
- Human gold labels for this title: the blocking step for any accuracy claim.
- Name/term prompts (cast, places) for the Whisper-family models.
- Ensemble BengaliAI + IndicConformer + Amazon Transcribe
bn-IN. - Managed comparators: Transcribe in the new AWS account, and Sarvam (whose code-mix mode keeps English in Latin) if a data agreement allows.
- Fine-tune the lead model on Hoichoi editorial subtitles, the most likely way to beat a general-purpose vendor.
- Word-level forced alignment for cue-timing parity.
Method and provenance
- Inputs
- 16 kHz mono WAV from the master; Silero VAD 6.2.3 (threshold 0.5, min speech 250 ms, min silence 500 ms, pad 200 ms, merge gaps ≤ 1 s, chunks ≤ 28 s).
- Decoding
- Language fixed to Bengali, temperature 0, no previous-text conditioning; greedy unless marked beam 5. Per-run settings are in each column’s details.
- Normalisation
- Every system was projected into one schema (start, end, text) with vendor-specific fields removed (romanization, speaker labels). Scores use NFC Unicode and strip punctuation.
- Windows
- A cue belongs to the speech chunk containing its midpoint (±0.5 s). Tessact cues outside every chunk get their own window.
- Reproduce
- Private run store
runs/poc-20260928: raw outputs, per-run manifests with checkpoint and input hashes, scorer (score.py --selftest), and all score files. Data for this page: asr-diff.json. - Not yet measured
- Accuracy against human labels, word-level timing for chunk-level models, GPU (AWS) throughput, Amazon Transcribe, Sarvam.