Picture this.

Imagine a recruiting team using AI to turn interview notes into a first draft of a candidate summary. The model was trained before the team ever opened the tool. When a recruiter submits those notes and gets a draft back, the model is performing inference.

Why it matters.

The work doesn’t end when a model has been trained. Every new request takes computing resources. Across a business, the speed, cost and reliability of those requests affect whether a tool is useful in everyday work.

The bit to remember.

A confident answer is still a generated answer. Inference does not mean the model has checked the facts. In our recruiting example, a person still needs to check the summary against the original notes.

“How will we check the output before somebody acts on it?”

Further reading: Cloudflare: inference and training. The business example above is illustrative.