PyTorch has two useful ways to run a block without recording operations for backward: torch.no_grad() and torch.inference_mode(). Their immediate effect looks similar. The difference appears when a tensor created inside the block reaches later training code. PyTorch’s autograd guide says outputs created under no-grad can be used in gradient-tracked computations later. Tensors created in inference mode carry a stricter limitation.
No-grad leaves the door open
No-grad mode treats computations in its block as though their inputs do not require gradients. Those computations stay out of the backward graph. Once the block ends, its output tensors can still participate in later operations that autograd records. PyTorch gives optimizer updates as an example: the update itself is untracked, while the updated parameters take part in the next tracked forward pass. The guide explains this distinction.
Inference tensors skip more bookkeeping
Inference mode also keeps its operations out of the backward graph. It additionally skips view tracking and version-counter updates, which reduces autograd overhead. Newly allocated tensors in that mode become inference tensors. PyTorch’s inference-mode reference describes the restriction: those tensors cannot be used in computations recorded by autograd.
That restriction follows from the bookkeeping inference mode omits. Some backward operations save tensors from the forward pass. PyTorch uses version counters to check whether a saved tensor changed before backward reads it. Inference tensors have no version counter. The autograd guide explains saved tensors and the check; PyTorch’s gradient-mode reference states that reading an inference tensor’s version raises an error.
What to do
Use torch.inference_mode() for work whose new tensors stay outside gradient-tracked computations, such as a self-contained evaluation path. If those tensors must flow into later tracked work, use torch.no_grad() for the producing block. PyTorch recommends that switch. If you already have an inference tensor and need a regular tensor, clone it after leaving inference mode; PyTorch documents that conversion. Set model.eval() separately when evaluation behavior matters: inference mode does not set it for you. The inference-mode reference makes that distinction explicit.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.