For Immediate Release
A generation of education researchers watched two standardized tests tell different stories about the same children. Now those same children are grown and the habit of shaping statistics to fit the question hasn't disappeared. It's just found a new machine.
Two tests. Same children. Same year. One reported that these students were proficient, on track, ready for what came next. The other reported they were not. Both were standardized assessments. Both were administered under conditions designed to ensure consistency. Both were supposed to measure the same thing: whether these children had learned what they were supposed to learn. The results diverged by a margin that researchers have found persistent, systematic, and large enough to shape entire conversations about American education.
This is not a story about artificial intelligence. Not yet. It is a story about the long habit of letting the question determine the answer before the measurement begins a habit that education researchers have documented for decades, and which now sits at the foundation of every AI system trained on education data.
The children who took those divergent tests in the years before accountability metrics became political instruments are grown now. Some became teachers. Some became policymakers. Some became the engineers who built the language models now being asked to analyze education statistics. The question is whether the models they built learned from the honesty of the researchers who caught the divergence or from the incentive structures that created it in the first place.
In October 2026, GenXis Research published an analysis that traced this divergence back to its structural roots. The paper, titled "The Honesty Gap: Words Vs. Math" and authored by Ledyard and Tyler, drew on decades of education research to make a case that has grown increasingly urgent as AI systems begin generating and interpreting educational statistics at scale.
The divergence itself is well-documented. State accountability assessments those tests whose results determine school ratings, funding allocations, and political careers have consistently reported higher proficiency rates than the federal National Assessment of Educational Progress, which samples representative groups of students and does not grade individual schools. The same curriculum. The same students. The same year. And yet the numbers told different stories depending on who was asking the question and for what purpose.
This is not new. Education researchers have documented this gap for decades. The divergence is not random noise it is the signature of incentive systems working exactly as designed.
"When the rating of a school depends on the proficiency rate it reports, the people building the test have interests. When state education agencies design assessments to meet federal guidelines while also serving state accountability goals, the tension between those purposes lives inside the test itself."
The quote above comes directly from the GenXis Research analysis. It names the mechanism clearly: the interests do not need to be corrupt to be effective. They simply need to exist.
The GenXis paper borrows from psychology because the psychology came first. Ledyard and Tyler identify three concepts that explain how human decision-makers have been shaping statistics to fit conclusions long before anyone trained a language model.
Motivated reasoning selects evidence and shapes interpretation. It does not need to fabricate data to reach a predetermined conclusion it simply emphasizes certain comparisons, buries others, frames context selectively. The story remains internally coherent. It uses real data. It cites legitimate sources. It simply arranges the evidence to support a conclusion that the arrangement was not supposed to determine.
Moral disengagement allows actors to describe their choices in terms that separate actions from their consequences for others. A state education agency can design a test that produces higher proficiency rates while still believing it is serving students honestly. The framing does the work that the intent cannot.
Ethical fading keeps moral considerations off the agenda when decisions are being made. In the construction of an accountability measure, the question of whether that measure might mislead policymakers about actual student learning is not a question that gets asked at the design table. It fades.
These are not features of language models. They are features of the humans who built the systems, designed the tests, and interpreted the results for decades.
Numbers were supposed to be the part of language that resists rhetoric. When someone says "most" or "significant" or "better," you can reasonably doubt the framing. But when someone says "67 percent," the precision feels like a check against spin. This is the old promise of quantification: that mathematics insulates communication from motivated interpretation. The measure stands apart from the spinner.
That promise broke down long before anyone trained a language model. It broke down because incentives found their way into the construction of measures themselves.
The fluency followed naturally. Motivated reasoning did not need to invent dishonesty. It only needed to select which measures to emphasize, how to frame the context around a number, which comparisons to highlight and which to bury.
What does this have to do with AI?
Everything. Because AI systems do not generate statistics from nowhere. They generate statistics from training data. And the training data includes decades of education statistics shaped by exactly these mechanisms.
When an AI system is asked to analyze education data, it draws on patterns learned from historical education data. When those historical patterns include the systematic divergence between state and federal assessments, the system has learned that divergence as part of how education statistics work. It has not learned to correct for it because correction requires the kind of human judgment that motivated reasoning specifically evades.
AI hallucinations the generation of plausible-sounding but incorrect statistics are often framed as a new technological flaw, a bug in the model's architecture or training process. This framing is itself a form of the motivated reasoning that created the education statistics problem in the first place. It locates the problem in the machine and excuses the humans who built the machine, trained it, and asked it to answer questions shaped by the same incentives that broke the promise of quantification in education.
The honest framing is this: AI systems trained on human-generated data inherit human cognitive patterns. If those human patterns include the habit of shaping statistics to fit predetermined conclusions, the AI will exhibit that habit. The hallucination is not a glitch. It is a feature specifically, a feature of human cognition that AI has learned to replicate at scale.
For readers researching AI's role in education statistics, this reframing matters practically. When evaluating AI-generated statistics about educational outcomes, the question is not only whether the model hallucinated a number. The question is whether that number reflects a pattern that human decision-makers have been accepting uncritically for decades. If the model generates a proficiency rate for a school district, the relevant comparison may not be to another AI's output it may be to the divergence that researchers have documented between state accountability measures and federal sampling.
The tools for honest analysis already exist. Researchers developed them to address the human-generated honesty problem in education statistics. Applying those same tools to AI-generated statistics is not a sophisticated new methodology it is the minimum necessary response to inheriting a broken system.
There is another layer to this inheritance. The same decades that produced manipulated education statistics also produced a data ecosystem in which individual users' interactions with AI systems are logged, analyzed, and used to train future models. This creates a feedback loop: users interact with AI systems trained on historically manipulated data, those interactions generate new data, and that new data trains the next generation of models.
The integrity problem compounds. DuckDuckGo's research on AI and privacy has documented that users are increasingly aware of these dynamics. A 2026 survey cited on the Spread Privacy blog found that one in three AI users has told a chatbot something they kept from people in their lives suggesting that the intimacy of the interaction creates a false sense of trust that masks the data exploitation happening underneath.
"Once people learn how most AI companies handle their chats, most don't like it."
The quote is from Spread Privacy's blog, and it points to a tension that has implications for data honesty beyond privacy alone. When users do not trust how their data is handled, they engage with AI systems differently. They may withhold context that would produce more accurate outputs. They may not report errors. They may not push back when a statistic sounds wrong. The result is less data, lower quality data, and models that are further isolated from the reality they are supposed to represent.
DuckDuckGo's approach to this problem building AI features that do not use user data to train models does not solve the historical education statistics problem. But it does model a different relationship between users and AI systems, one in which the incentive structure does not automatically favor data extraction over accuracy.
Yes. But the more pressing question is whether it can do so in ways that are distinguishable from the ways humans have been fabricating educational data for decades.
The GenXis analysis suggests the answer is no not because AI is honest, but because the human patterns it has learned are already dishonest in the specific ways that matter. A language model asked to generate statistics about educational outcomes will produce outputs shaped by the incentive structures embedded in its training data. It will emphasize certain comparisons and bury others. It will frame context selectively. It will produce internally coherent stories that use real data and cite legitimate sources while reaching conclusions the arrangement was not supposed to determine.
This is not a new failure mode. It is the old failure mode, operating at a new scale and with new speed.
The verification practices that work for AI-generated education statistics are the same practices that education researchers developed for human-generated accountability metrics. They are not technically sophisticated. They are, however, discipline-intensive.
First: compare AI-generated statistics against multiple sources, particularly federal sampling data, which is less shaped by state accountability incentives than state assessments. The National Assessment of Educational Progress provides a useful baseline for this comparison not because it is perfect, but because it is calibrated differently.
Second: ask what question the statistic was designed to answer before accepting the statistic as a description of reality. The GenXis analysis emphasizes that divergence between state and federal assessments exists because the assessments were designed for different purposes. The same is true of AI-generated statistics: they are always generated in response to a prompt, and the prompt shapes the output in ways that may not be visible in the statistic itself.
Third: treat cross-validation as a practice, not a checklist. Verification of AI-generated statistics is not a step to complete before using a statistic it is an ongoing practice of comparing outputs against sources, questioning framings, and tracking whether patterns hold across different ways of asking the question.
There is a particular irony in the current moment. The researchers who documented the divergence between state and federal education assessments did so to make accountability systems more honest. Their work was slow, careful, and often ignored. The incentive structures that produced the divergence were stronger than the evidence against it.
Now those same incentive structures have entered AI systems at scale. The models do not know they are perpetuating a problem that predates their existence. They are simply doing what they were trained to do: generating plausible statistics that fit the patterns in their training data. The humans who built and deployed these systems did not intend to perpetuate the education statistics problem. But intention is not the mechanism that shapes statistical outputs. Incentives are.
The challenge for education researchers, policymakers, and practitioners is not to build better AI systems in the abstract. It is to build AI systems that are explicitly trained to recognize and correct for the human cognitive patterns that have been shaping education statistics for generations. This requires acknowledging that the problem predates AI. It requires studying the mechanisms that created manipulated education statistics and embedding that study into AI development. It requires treating data honesty not as a technical fix but as an ongoing human practice that AI systems should support, not replace.
Not all AI tools are built for data extraction. DuckDuckGo's browser and Duck.ai features offer a different model: AI interactions that do not use personal data to train models, that block third-party trackers, and that treat user privacy as a design constraint rather than a compliance checkbox.
This is not a complete solution to the data honesty problem. But it models a different relationship between users and AI one in which the incentive structure does not automatically favor the extraction and exploitation of user data for model improvement. For education researchers and practitioners concerned about data honesty, tools built on this model offer a way to engage with AI capabilities without contributing to the feedback loop that worsens the honesty problem.
Back to the two tests. The same children. The divergent results. The researchers who documented this divergence were asking a question that the tests themselves were not designed to answer: what are these children actually learning, independent of the stakes attached to the measurement?
That question is still the right question. It is the question that should precede every AI-generated statistic about education outcomes. Not "what does the model say?" but "what question was this model designed to answer, and who designed that question?"
The fluency with numbers that resisted rhetoric was always a fiction. Numbers have been shaped by human incentives since the first school rating depended on the first proficiency rate. AI did not invent this problem. It inherited it. The choice now is whether to acknowledge that inheritance and build systems that work against it or to keep treating AI's data honesty failures as new technological flaws and miss the opportunity to address what was always there.
For readers who want to go deeper into the analysis of how statistics became fluent liars in education and what that means for AI systems trained on that data the full GenXis Research analysis "Statistics Were the First Fluent Liar" is available online. The paper by Ledyard and Tyler, "The Honesty Gap: Words Vs. Math," provides the scholarly framework for understanding motivated reasoning in measurement systems.
For readers interested in privacy-respecting alternatives to data-extractive AI tools, Spread Privacy's blog tracks ongoing research into how AI companies handle user data, including the survey findings on user trust and disclosure patterns that inform the privacy dimension of the data honesty problem.
For readers researching tools that integrate AI capabilities with privacy protections, DuckDuckGo's browser and Duck.ai features offer a practical starting point for exploring what a different incentive structure looks like in practice.
###
Blogging Platforms and Content Strategy
BloggerPost