GenAIWiki
Serving

Inference

Inference is running a trained model on new input to produce an output. Training updates the model's weights. Inference does not.

Expanded definition

A model is trained or fine-tuned, then served. Each API call, chat reply, image, or transcription is inference. Latency, price per token, and rate limits describe inference, not training. Batch and cached prompts are still inference; they change cost and speed, not whether the weights are being updated.

Related terms

Explore adjacent ideas in the knowledge graph.

Inference FAQ

What is Inference?

Inference is running a trained model on new input to produce an output. Training updates the model's weights. Inference does not.

How is Inference used in AI systems?

A model is trained or fine-tuned, then served. Each API call, chat reply, image, or transcription is inference. Latency, price per token, and rate limits describe inference, not training. Batch and cached prompts are still inference; they change cost and speed, not whether the weights are being updated.

Related

Comparisons, tools, and models that connect to this idea.