f316b9bbbc
IT狗今日升級嘅 transcribe server: - Model: large-v3-turbo → large-v3 (+2-3% Cantonese WER improvement) - Compute: cpu + int8 (MPS fp16 unstable for large-v3 on 16GB Mac) - Beam: 5 → 1 (greedy, ~30-40% speedup) - Added: condition_on_previous_text=False (anti-hallucination chain) - Added: compression_ratio_threshold=2.4 (repetitive noise filter) - Added: no_speech_threshold=0.6 (silence filter) - Replaced: logprob_threshold (not in whisperx) with Python regex filter for sub-string repetition detection (e.g. '对,有自己的监控,IP' × 20) - Test: 6.1MB / 6:38 voice memo → 358s = 1.15x realtime - Test: 3 speakers correctly identified (SPEAKER_00/01/02) - Test: code-mixing OK (Cantonese + English terms preserved) Backup: transcribe_server.py.bak-20260724-0931
9 lines
74 B
Plaintext
9 lines
74 B
Plaintext
# Backup files
|
|
*.bak.*
|
|
*.orig
|
|
# Mac
|
|
.DS_Store
|
|
# Python
|
|
__pycache__/
|
|
*.pyc
|