Interestana
Home/News/LLMs Corrupt Documents in Editing, Microsoft Research Finds
Neil Patel5 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

LLMs Corrupt Documents in Editing, Microsoft Research Finds

LLMs Corrupt Documents in Editing, Microsoft Research Finds

A Microsoft Research study published on April 17, 2026, has revealed that large language models (LLMs) can subtly and systematically corrupt documents during extended editing workflows. The DELEGATE-52 study tested 19 LLMs across 52 professional domains over 20 editing interactions, simulating real-world document work. The research found that even advanced frontier LLMs, including Gemini 3.1 Pro, Claude 4.6 Opus, and GPT 5.4, corrupted an average of 25 percent of document content by the end of these long editing sessions. Across all 19 LLMs tested, the average degradation reached a significant 50 percent. The errors introduced by LLMs are described as sparse but severe, consisting of a small number of consequential changes that are grammatically correct, making them difficult to detect during a casual review. This contrasts with typical hallucinations or minor typos, as these errors are subtle enough to pass initial scrutiny but damaging enough to impact the document's integrity and can compound over multiple editing sessions. The study's methodology involved providing LLMs with professional documents from diverse fields such as coding, crystallography, music notation, accounting records, and recipes. These domains encompassed both highly structured formats like code and database schemas, and natural language writing, demonstrating that corruption occurred across various types of content. The DELEGATE-52 study specifically focused on multi-session workflows, where an LLM handles a continuous sequence of revisions and refinements, rather than isolated, one-off edits. This approach aimed to mimic how users might employ LLMs for ongoing document development. The researchers also investigated the impact of providing LLMs with basic agentic harnesses and file tools. Counterintuitively, this slightly worsened performance, leading to approximately 6 percent more degradation while consuming two to five times more input tokens. Python emerged as the only domain where most LLMs managed to clear the study's stringent 98 percent accuracy threshold. However, even the best-performing model achieved this benchmark in only 11 out of the 52 domains tested, highlighting the widespread nature of the corruption issue. These findings suggest a need for content teams to re-evaluate the integration of AI tools into their editing processes, particularly for tasks involving iterative document refinement.

Original source — read the full reporting at the publisher:

Read on Neil Patel

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next