Fluent and Wrong: Eight Failure Patterns in Extended AI Use
Research on large language model failure predominantly assesses models in single exchanges and isolation: a model is prompted, the output is scored, and an error rate is reported. That is not how these systems are used in professional practice. Clinicians, attorneys, analysts, educators, and researchers work with them across extended sessions in which each output becomes the context for the next, and the final product carries a human signature. A test that scores a single response will not reveal these emerging behaviors. This paper describes eight failure patterns observed across extended sessions on multiple commercial AI platforms: recognition without correction, compounding rather than isolation of errors, halo effect from real anchors, unreliable self-explanation, degradation across long sessions, motivated framing under confrontation, sycophancy and confirmation bias amplification, and an architectural rather than statistical failure profile. Several are consistent with findings already established in the machine learning literature. Holistically, they are not addressed by lower hallucination rates, larger models, or vendor-side mitigations and instead are properties of system behavior under extended engagement, which carry direct consequences for verifying AI-assisted work in any domain where a person signs what the machine produced.