Definition

LLM reliability is the degree to which large language model outputs remain factually grounded, consistent, and appropriate for a given decision context — including known failure modes like hallucination, sycophancy, and format mimicry.

Key Points

  • Military intelligence near-miss shows LLM summaries can inherit official report formatting without independent verification (ai-hallucination)
  • Alternative architectures (e.g. jev System One models) structurally constrain outputs to predefined tokens with calibrated probabilities — trading generality for reliability on classification tasks
  • Human-in-the-loop and corroboration requirements scale with decision stakes

Sources