← Tessact
Benchmark · Bengali ASR

Tessact vs open models: Bengali transcription

One Hoichoi episode, transcribed by Tessact twice and by twelve open-weight configurations on identical speech chunks. Compare them window by window.

Calibration title · Meyebela S1E1 Laptop run · $0 cloud spend

Tessact is a reference, not ground truth. Every score here measures agreement with Tessact’s 20 Sep output. Where a model and Tessact disagree, either may be right. Human gold labels (runbook Phase 2) are needed before anyone can say a system is better than Tessact.

The title

Title
Meyebela, Season 1, Episode 1: “Roommates forever?” (Bengali drama; Kolkata flat-share, colloquial Bengali with English code-mix). Watch on hoichoi.tv/shows/meyebela.
Media
762.0 s (12 min 42 s) · H.264 1920×1080 25 fps, AAC · 1.87 GB master
Fingerprint
SHA-256 fe695db1…3603; the same bytes Tessact processed on 20 Sep and again in the controlled reingest on 27/28 Sep
Speech
VAD found 87 speech regions, merged into 43 chunks totalling 300.5 s. Every open model transcribed exactly these chunks, so only the model differs.
Hardware
Apple M5 Pro, 64 GB. Speeds are laptop speeds; the AWS GPU run is still to come.

Scores

Three views, because whole-file comparison is dominated by conventions rather than recognition (see why). CER and WER are against Tessact 20 Sep; lower means closer to Tessact.

ConfigurationGroup× real timeCER · whole fileCER · chunk, Bengali scriptCER · pure-Bengali chunksWER · pure-BengaliLicence
Loading scores…

Pure-Bengali view: the 8 speech chunks whose Tessact text has no Latin letters (680 characters). Small sample; indicative only.

Diff viewer

Each row is one speech chunk, or a caption some system placed outside detected speech (italic). Pick a base column; every other column highlights words it adds relative to the base (added) and, optionally, base words it misses (missed). Collapse a column with −; open details in a header for run metadata. The URL keeps your selection, so a link opens the same view.

Why whole-file scores mislead

Non-speech captions
14 of Tessact’s 92 cues (444 characters) sit mostly outside detected speech, including a 65 s cue over the music intro and a 139 s “Good morning.” cue. Filter captions outside speech to see them.
Code-mix script
23% of Tessact’s characters are Latin: English words written in English. Every open model writes them phonetically in Bengali script (“Time bound” → “টাইম বাউন্ড”). That counts as an error without being a mishearing. Tick Bengali script only to hide Latin text.
Uncaptioned English
Some chunks contain English speech Tessact did not caption at all while the models did (e.g. a 10 s line where Tessact has only “মানে মানে”). These count against the models but may be Tessact omissions.
Spelling
Colloquial variants such as পারবোনি / পারবো নি, আনব / আনবো, দাও / দেও differ in characters, not meaning.

What the diff shows

Lead model
BengaliAI regional Whisper-medium (Apache-2.0) is closest to Tessact on pure Bengali (CER 0.17 with beam 5, 0.19 greedy) and never loops. On clean dialogue it is often near-identical; one 12 s chunk differs by 2%, in spelling only.
Fine-tunes win
Bengali fine-tuned medium-size models beat multilingual Whisper large-v3 by a wide margin at a fraction of the compute.
Fast second opinion
IndicConformer 600M runs ~120–140× real time on CPU with mid-table accuracy, which makes it useful for ensembling or a cheap first pass.
Rejected
Whisper large-v3-turbo loops on this audio (21–23 looping segments). Whisper large-v3 without VAD hallucinates repeated words over music.
Real errors
Place names (উলুবড়িয়া → কুলবাড়িয়া) and very short, noisy chunks.
Vendor noise floor
Tessact vs its own reingest of identical bytes: CER 0.04–0.06, scene count 6 → 8. Any system scoring near that is within the vendor’s own run-to-run variation.

Path to match or beat Tessact

  1. Fix the code-mix convention (English in Latin vs Bengali script) and add a transliteration pass.
  2. Spelling-tolerant scoring, so recognition is measured rather than orthography.
  3. Human gold labels for this title: the blocking step for any accuracy claim.
  4. Name/term prompts (cast, places) for the Whisper-family models.
  5. Ensemble BengaliAI + IndicConformer + Amazon Transcribe bn-IN.
  6. Managed comparators: Transcribe in the new AWS account, and Sarvam (whose code-mix mode keeps English in Latin) if a data agreement allows.
  7. Fine-tune the lead model on Hoichoi editorial subtitles, the most likely way to beat a general-purpose vendor.
  8. Word-level forced alignment for cue-timing parity.

Method and provenance

Inputs
16 kHz mono WAV from the master; Silero VAD 6.2.3 (threshold 0.5, min speech 250 ms, min silence 500 ms, pad 200 ms, merge gaps ≤ 1 s, chunks ≤ 28 s).
Decoding
Language fixed to Bengali, temperature 0, no previous-text conditioning; greedy unless marked beam 5. Per-run settings are in each column’s details.
Normalisation
Every system was projected into one schema (start, end, text) with vendor-specific fields removed (romanization, speaker labels). Scores use NFC Unicode and strip punctuation.
Windows
A cue belongs to the speech chunk containing its midpoint (±0.5 s). Tessact cues outside every chunk get their own window.
Reproduce
Private run store runs/poc-20260928: raw outputs, per-run manifests with checkpoint and input hashes, scorer (score.py --selftest), and all score files. Data for this page: asr-diff.json.
Not yet measured
Accuracy against human labels, word-level timing for chunk-level models, GPU (AWS) throughput, Amazon Transcribe, Sarvam.