Inference & Serving
25 min
Structured Generation and Constrained Decoding: Making LLMs Predictable
Language models generate text one token at a time by sampling from a probability distribution over their entire vocabulary. Constrained decoding intervenes at that sampling step, masking out every token that would violate a targe…