IT狗 335b0d8ee5 FIX: pre-encode mp3→wav + preloaded audio dict for pyannote (FFmpeg 8 / torchcodec 0.7 broken)
ROOT CAUSE
- whisperx venv has torchcodec 0.7.0 + PyTorch 2.8.0 → torchcodec dylibs target
  FFmpeg 4-7 (libavutil.56-.59) but Mac brew now ships FFmpeg 8 (libavutil.60).
- All audio loading via torchcodec silently fails (warning but hangs/returns 0).

FIX
- /transcribe endpoint: pre-encode uploaded mp3/webm → 16kHz mono WAV via
  subprocess ffmpeg (whisperx.load_audio then reads clean PCM file).
- _do_transcribe diarize branch: pass {'waveform': tensor, 'sample_rate': int}
  dict to pyannote Pipeline instead of file path → skips torchcodec entirely.

Tested
- chunk_0.mp3 (4MB, 5min): diarize=0 → 200 OK + hallucination filter triggered
- Confirmed working in live HTTP POST at 15:40 HKT
2026-08-11 15:47:00 +08:00

Transcribe Server (M1 Mac)

Local Whisper transcription server with speaker diarization.

Stack

  • Model: large-v3 (int8, CPU) — best quality for Cantonese
  • Beam size: 1 (greedy, fast)
  • Diarization: pyannote/speaker-diarization-3.1
  • Framework: FastAPI + uvicorn
  • Port: 8765 (local) → 18765 (via SSH reverse tunnel from VPS)

Why large-v3

  • Cantonese WER improvement ~2-3% over medium / large-v3-turbo
  • Better English code-mixing preservation
  • More accurate speaker diarization
  • Trade-off: ~4x slower, ~1GB more RAM

v1.3.0 (2026-07-24)

  • Upgraded from large-v3-turbo → large-v3
  • Added hallucination filter (repetition collapse + segment dropping)
  • Replaced inline logprob_threshold with Python-level regex filter
  • Faster: 1.15x realtime for Cantonese voice memo (vs ~1.5x for large-v3-turbo)

Usage

# local
python3 transcribe_server.py

# test
curl -X POST -F "file=@/path/audio.m4a" \
  -F "language=cantonese" \
  "http://127.0.0.1:8765/transcribe?diarize=1" \
  -o output.json

Hallucination Filter (v1.3.0+)

  • Detects raw text repetition (e.g. "对,有自己的监控,IP" × 20)
  • Drops segments with high compression_ratio (>2.4)
  • Drops segments with no_speech_prob > 0.6
  • Replaces broken logprob_threshold (not in whisperx TranscriptionOptions)

Auto-restart

Managed by launchd: ~/Library/LaunchAgents/com.itdog.transcribe-server.plist

  • Restarts on crash
  • Loads model on first request (lazy)

Integration

VPS meeting-bot at https://meet.donton.cloud/upload calls transcribe_server via SSH reverse tunnel:

  • VPS port 18765 → Mac port 8765
  • See /opt/meeting-bot/backend/main.py for backend
S
Description
Local whisper transcribe server (M1 Mac) - large-v3 + hallucination filter
Readme 47 KiB
Languages
Python 100%