9a93cb6ffde0444ab48e11d43cde520282d682a8
Speaker labels were empty because whisperx 3.8.5's DiarizationPipeline wrapper has a compatibility bug with pyannote.audio 4.0.4: the wrapper silently returns empty segments (the 'fallback to plain text' message in 3.8.5 logs), so segments.speaker becomes '' instead of 'SPEAKER_XX'. Fix: load pyannote.audio.Pipeline directly and convert DiarizeOutput → DataFrame the same way whisperx/diarize.py does internally. Verified with 30s Cantonese clip from real recording (chunk_0.mp3): speaker: SPEAKER_00 (was '') text: '那就係F5嘅Low Balancer用喺你哋嘅...特別感謝Sherry...' took: 34s for 30s audio (1.13x realtime, CPU int8 + diarize=1) curl http://localhost:8765/transcribe -F file=@test.wav -F language=yue -F diarize=1
Transcribe Server (M1 Mac)
Local Whisper transcription server with speaker diarization.
Stack
- Model: large-v3 (int8, CPU) — best quality for Cantonese
- Beam size: 1 (greedy, fast)
- Diarization: pyannote/speaker-diarization-3.1
- Framework: FastAPI + uvicorn
- Port: 8765 (local) → 18765 (via SSH reverse tunnel from VPS)
Why large-v3
- Cantonese WER improvement ~2-3% over medium / large-v3-turbo
- Better English code-mixing preservation
- More accurate speaker diarization
- Trade-off: ~4x slower, ~1GB more RAM
v1.3.0 (2026-07-24)
- Upgraded from large-v3-turbo → large-v3
- Added hallucination filter (repetition collapse + segment dropping)
- Replaced inline
logprob_thresholdwith Python-level regex filter - Faster: 1.15x realtime for Cantonese voice memo (vs ~1.5x for large-v3-turbo)
Usage
# local
python3 transcribe_server.py
# test
curl -X POST -F "file=@/path/audio.m4a" \
-F "language=cantonese" \
"http://127.0.0.1:8765/transcribe?diarize=1" \
-o output.json
Hallucination Filter (v1.3.0+)
- Detects raw text repetition (e.g. "对,有自己的监控,IP" × 20)
- Drops segments with high compression_ratio (>2.4)
- Drops segments with no_speech_prob > 0.6
- Replaces broken
logprob_threshold(not in whisperx TranscriptionOptions)
Auto-restart
Managed by launchd: ~/Library/LaunchAgents/com.itdog.transcribe-server.plist
- Restarts on crash
- Loads model on first request (lazy)
Integration
VPS meeting-bot at https://meet.donton.cloud/upload calls transcribe_server via SSH reverse tunnel:
- VPS port 18765 → Mac port 8765
- See
/opt/meeting-bot/backend/main.pyfor backend
Description
Languages
Python
100%