Does temperature 0 make an LLM deterministic?
hardAnswer
- Not in practice.
- Temperature 0 makes the sampling step greedy, but it does not remove every source of variation.
- Floating-point non-determinism in batched GPU kernels means the logits themselves can differ slightly between runs, and ties can flip.
- Requests are batched together differently depending on concurrent traffic, which changes reduction order.
- Mixture-of-experts routing can vary with batch composition.
- Providers also update model versions behind a stable alias.
- Treat low temperature as reduced variance, not a guarantee, and if you need reproducibility, pin the model version and store the outputs you depend on.
Check yourself — multiple choice
- Yes, it is fully deterministic
- No — greedy sampling still sits on non-deterministic batched float arithmetic, variable batching, MoE routing and silent version updates
- Only if you set a seed
- Temperature has no effect
Greedy decoding reduces variance but does not guarantee bit-identical output.
#decoding#sampling#production
Practise LLMs & GenAI
214 interview questions in this topic.