A short answer may end at a tidy sentence or halfway through one. Its shape alone does not tell you why generation stopped. Hugging Face Transformers exposes separate controls for stop strings, token length, and run time. Check the settings before treating every cut-off as the same problem.
A stop string matches generated text
stop_strings is a string or list of strings that ends generation when the model outputs one of them. This is useful when a response has a deliberate delimiter. It can also end a response earlier than expected if the same text appears in ordinary output.
For example, if an application uses END as its stop string, an answer that reaches END has met that condition. Raising the token allowance leaves the string condition in place. Inspect the configured strings and the generated text together. Hugging Face describes this condition.
A token cap limits new output
max_new_tokens sets the maximum number of generated tokens and ignores tokens in the prompt. A response that reaches that allowance can stop mid-thought even if no stop string appears. Transformers also lists max_length, but recommends max_new_tokens for controlling generated length. The generation configuration explains both fields.
Compare the number of new tokens with the configured cap. Use the token count rather than the visible word count. If the cap is the likely cause, increase it only as much as the task needs, then check the result again.
A time limit lets the current pass finish
max_time sets a time allowance in seconds. Transformers says generation still finishes its current pass after that allowance expires. The time limit works independently of a token count or a matched string. See the documented time behavior.
If runs stop around the configured time while using fewer than the allowed new tokens, inspect the time setting. Increasing the token cap alone does not extend that allowance.
What to do
Record the effective generation settings for the call. In Transformers, a supplied generation configuration forms the base, and matching arguments passed to generate() override it. Otherwise, model defaults can supply values. Check the configuration rules.
Then compare the output with each setting: look for a configured stop string, count newly generated tokens, and check elapsed time against max_time. Change one limit at a time and repeat the same request. If none explains the ending, inspect other documented stopping conditions, including the end-of-sequence token and custom stopping criteria. Transformers documents both.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.