How I Made CosyVoice3 2.2× Faster on Apple Silicon
How I bounded PyTorch MPS memory growth, cut CosyVoice3 latency below real time, and moved only the dispatch-bound LLM stage to MLX.

Tag archive
Notes connected by the Text-to-Speech tag.
How I bounded PyTorch MPS memory growth, cut CosyVoice3 latency below real time, and moved only the dispatch-bound LLM stage to MLX.
A practical quality loop for local text-to-speech: representative fixtures, contextual normalization, useful latency metrics, defect labels, and release checks.
How to choose, benchmark, package, and isolate a local text-to-speech stack without treating a good demo as production evidence.
No entries are filed here yet.
Browse all notes →