Prefill vs decode explained: why one machine processes prompts at 1,700 tokens per second and still generates at 38, and ...