Why Models Look Great on Summarization but Fail at Knowledge Reliability
https://hannahsinsightfulnews.theburnward.com/what-makes-claude-better-than-gpt-at-finding-hidden-assumptions
Benchmarking the Disconnect Between Low Vectara High AA Omni Hall and Actual Knowledge Reliability Decoding the Summarization Performance Gap As of March 2026, I have spent a significant amount of time staring at the divergence