Where it stands. Tessact produces rich, timecoded Bengali metadata, but its output changes between runs of the same file, it has no supported export API, and its price is unknown. An open Bengali speech model already gets close to Tessact’s transcripts on clean dialogue. Accuracy claims wait on human-labelled ground truth.
Pages
- 01Product evidence and build optionsWhat Tessact does, its private API and data model, integration gaps, a controlled re-upload, and the full evidence export.done
- 02Bengali transcription benchmarkTessact vs 12 open models on one episode, with a side-by-side diff viewer.done
- 03Managed speech APIsAmazon Transcribe (Bengali) and others, added as columns to page 02.awaiting approval
- 04GPU speed and cost on AWSThroughput and cost per media hour on one cloud GPU.awaiting approval
- 05Scenes, shots and visual descriptionsScene boundaries and scene descriptions against Tessact.planned
- 06Human-labelled accuracyIndependent labels that turn agreement scores into accuracy.needs annotators
- 07Cost and recommendationAPI vs own GPU vs hybrid, and whether to run a 15-title pilot.planned
Reference video
- Title
- Meyebela, Season 1, Episode 1: “Roommates forever?” (12 min). Four friends share a Kolkata flat; a renewal notice upends a chaotic morning.
- Watch
- hoichoi.tv/shows/meyebela (episode 1). Subscriber content.
- Used for
- Every benchmark so far: Tessact processed it on 20 Sep and again from identical bytes on 27/28 Sep; all open models ran on the same audio.
Key numbers
| What | Value | Note |
|---|---|---|
| Titles Tessact has analysed in our workspace | 8 | 7 video, 1 audio |
| Upload → finished analysis, 12.7-min file | 39.6 min | One run; 28.9 min after upload completed |
| Same file, one week apart: scenes | 6 → 8 | Tessact output is not stable between runs |
| Tessact vs itself, transcript difference | 4–6% | Character-level; the noise floor for any comparison |
| Best open model vs Tessact, pure-Bengali speech | 17% | BengaliAI regional Whisper-medium; not yet checked against human labels |
| Estimated cloud cost, 15-title pilot | $35–110 | Human labelling (90–180 h) dominates |
| Tessact price | unknown | No credits deducted in our test; no quote yet |
Status and next steps
- Done
- Live product survey, export of all results, controlled re-upload, pricing and quota checks, PoC plan (independently reviewed), local Bengali speech benchmark.
- Waiting on
- Approval for small, spend-capped cloud tests; annotators for ground-truth labels; a decision on how English words inside Bengali dialogue should be written; a price quote from Tessact.
- Next
- Managed speech APIs, GPU throughput, scene and visual comparison, then a go/no-go on a 15-title pilot.