How I Made CosyVoice3 2.2× Faster on Apple Silicon
How I bounded PyTorch MPS memory growth, cut CosyVoice3 latency below real time, and moved only the dispatch-bound LLM stage to MLX.

Category archive
Field notes collected under Performance.
How I bounded PyTorch MPS memory growth, cut CosyVoice3 latency below real time, and moved only the dispatch-bound LLM stage to MLX.
A service token lived for more than ten hours, but our Gravitee gateway fetched it on every request. Here is how we measured the delay, tested the obvious cache fix, and found …
In this short piece, I want to share how you can make a website that performs well for users. Speed is very essential.
No entries are filed here yet.
Browse all notes →