PwC Completes the Big Four Set of AI Hallucination Failures
Every Big Four brand has now published reports containing fabricated or unsupported material, exposing a widening gap between the pressure to adopt AI and the controls needed to verify its output
Key Takeaways
The most troubling claims concerned governments supposedly using a PwC framework without supporting evidence.
The firms are pushing AI adoption at every level, from junior staff to partners.
Their traditional review model is poorly equipped to catch polished, plausible AI errors.
AI verification controls appear to be developing more slowly than adoption targets.
PwC has become the final member of the Big Four to be caught distributing reports containing serious apparent AI hallucinations.
An investigation published by GPTZero on 28 July 2026 identified fabricated citations, unsupported claims and misattributed evidence across four PwC Middle East thought-leadership reports published between 2024 and 2026. The findings were then verified by the Financial Times.
The reports covered government services, agentic AI, electric mobility and cybersecurity. GPTZero’s detector classified between 50% and 100% of each document as AI-generated or mixed, depending on the report. The use of AI is not itself the problem; the Big Four have been open about using it to improve productivity. The problem is that some of the claims and references were fabricated, unsupported or misattributed.
The most serious problem went beyond defective references. One of the reports, Transforming Governance: Citizen Pulse, a 2025 report promoting a PwC framework that purportedly uses real-time data, predictive analytics and artificial intelligence to improve government services, implied that public institutions in Denmark, Saudi Arabia, the United States and Australia were using the framework. Yet the sources cited did not substantiate those claims, and GPTZero said it found no corroborating public evidence that the particular framework described in the report had been deployed in those settings.
The three other reports contained further serious errors. One confused the number of electric cars sold in China since 2010 with the number of public charging points in the country. Another cited a non-existent academic paper about electric-vehicle adoption and air quality in Riyadh. A report on agentic AI attributed the use of NVIDIA technology to the Mayo Clinic before the relevant collaboration had been announced.
The Financial Times found that another report used a Medium post by a teenage blogger to support a JPMorgan “real-world success story”.
PwC Middle East said it took the accuracy of its research seriously and was updating a limited number of supporting citations. It said employees were expected to follow the firm’s quality-control processes for research and content development.
The Big Four set is now complete
The importance of the PwC case lies partly in what came before it.
Deloitte, EY and KPMG had already been linked to publications containing invented references, distorted sources or case studies that could not be substantiated. PwC’s inclusion means that all four global Big Four brands have now experienced a publicly documented failure of this kind.
The cases differ in seriousness. Some involved public thought leadership; others involved work commissioned by governments. Some reports were withdrawn, while others were revised. Only one led to a publicly confirmed refund.
They should not be treated as identical. But they reveal the same underlying problem: professional work carrying the authority of a Big Four firm reached its intended audience without the sources and claims having been checked effectively.
Deloitte Australia
Deloitte Australia’s Targeted Compliance Framework Assurance Review was commissioned by the Australian Department of Employment and Workplace Relations to examine the IT system used to administer automated penalties affecting welfare recipients. This was a deliverable with serious ramifications for people’s social security payments.
The report contained references to non-existent academic publications, distorted legal sources and a fabricated quotation attributed to a Federal Court judgment. The problems were first identified by University of Sydney researcher Christopher Rudge, who recognised that a book attributed to one of his colleagues did not exist.
An internal departmental review subsequently identified more than 50 suspect citations. GPTZero examined the report’s 141 references and identified more than 30 citation problems, including around 20 that it classified as probable hallucinations.
The report was revised and republished. Its methodology was updated to record the agreed use of a department-hosted generative-AI tool chain in the technical workstream. Departmental correspondence also recorded the use of generative AI to summarise a legal case and to complete and format citations, and said those tools likely contributed to the errors. Incorrect references and the erroneous account of the Federal Court proceeding were revised. Deloitte agreed to repay the final A$98,000 instalment of its A$440,000 contract.
Deloitte Canada
A separate controversy involved Deloitte Canada’s 526-page Health Human Resources Plan, prepared for the government of Newfoundland and Labrador at a reported cost of C$1.6 million.
The plan was intended to guide long-term decisions about healthcare recruitment, retention incentives, staffing shortages and virtual care. An investigation by The Independent found at least four citations to papers that did not appear to exist, including references that attached the names of real researchers to studies they said they had never written.
Deloitte later acknowledged that AI had been used “to support a small number of research citations”. It said it was correcting those citations but continued to stand behind the report’s conclusions and recommendations. Newfoundland and Labrador’s accounting regulator subsequently opened an investigation.
EY Canada
In May 2026, GPTZero examined a 44-page EY Canada cybersecurity report entitled Points of Attack: Uncovering Cyber Threats and Fraud in Loyalty Systems.
The report promoted EY’s expertise in protecting loyalty programmes from cybercrime and fraud. GPTZero found that almost all the URLs in its 27-entry resources table were broken, fabricated or misleading, while more than half the titles did not correspond to identifiable sources.
Examples included a purported Gartner report that did not exist, a fabricated McKinsey “Loyalty Economics Report”, a non-existent Wired article and several URLs that led nowhere. GPTZero also identified inaccurate statistics and claims that were not supported by the references provided.
EY removed the report from its website and said it was reviewing how it had been published. The firm emphasised that the document was not connected to a client engagement.
KPMG International
KPMG International followed in June 2026 when GPTZero examined its October 2025 report Total Experience: Redefining Excellence in the Age of Agentic AI.
Of the report’s 45 citations, only five accurately pointed to the sources described.
The report also presented supposed examples of agentic-AI adoption at organisations including UBS, Swiss Federal Railways, Transport for London and the UK National Health Service. Several organisations disputed how their activities had been described. KPMG withdrew the report and began an internal review.
The pressure to adopt AI runs from the bottom to the top
It would be easy to describe these incidents as isolated cases of careless employees using generative AI and failing to check the output.
That explanation becomes less credible as the failures accumulate across different firms, countries, service lines and levels of commercial importance.
The Big Four are actively encouraging their employees to use AI to increase output, reduce the time spent on routine work and redesign how professional services are delivered. In the US, KPMG has introduced a dashboard for approximately 10,000 employees in its advisory division that tracks how frequently they use AI and compares their usage with individual targets and peer-group benchmarks.
The pressure does not apply only to junior employees expected to complete more work in less time. It extends to the top of the partnership.
In March 2026, PwC US senior partner and chief executive Paul Griggs warned that partners who resisted AI would have no future at the firm. He said senior personnel who were not “paranoid about being AI-first” were likely to be replaced and that an employee who believed they could opt out was “not going to be here that long”.
Griggs made the comments as PwC was developing automated services and considering alternatives to the traditional practice of billing clients according to the hours worked by professional staff. His remarks were not simply an invitation to learn about a useful new tool. They framed AI adoption as a condition of continued relevance inside the firm.
However, beyond the hype about the billions being invested in new AI tools and claims that AI will transform the firms, we have heard little about the controls the firms have put in place to ensure output quality.
The pressure to adopt AI is visible, measurable and increasingly relevant to career progression. The controls intended to assure the accuracy of the resulting work are either less concrete, less consistently applied or insufficiently effective.
Verification consumes part of the efficiency gain
If AI is introduced chiefly as a way to save time, thorough verification consumes part of that gain. Someone still has to open each source, confirm that it exists, compare it with the claim and resolve any discrepancy. Those steps take time.
The organisational message is therefore mixed. Employees are told that AI is essential to their future, that it should be used frequently, and that it will allow the firms to deliver more work with fewer routine hours. They are also told that humans remain responsible for catching every fabricated source, distorted statistic and invented case study.
The first message is reinforced through leadership warnings, utilisation targets, investment commitments, internal platforms, awards and changes to pricing models. The second is usually expressed through broad principles such as responsible use, professional judgement and human oversight. The gap between the two is left for individuals to navigate.
The leverage model was not designed for AI
The main difficulty is the way large professional-services firms organise work.
The Big Four operate through a leveraged hierarchy. Junior staff perform much of the research, testing, analysis, data preparation and initial drafting. Managers and senior managers supervise the work and review the output. Partners manage the engagement or publication, resolve major questions and accept ultimate responsibility.
This structure is central to the firms’ economics. Partners cannot personally perform or reperform every procedure carried out by large teams. They rely on work passing through successive layers of preparation and review.
A partner relies on a senior manager’s assessment that the important issues have been addressed. The senior manager relies on the manager’s review. The manager relies on the staff who performed the work and assembled the evidence.
That model was developed around human preparation and familiar forms of human error. A junior employee might use the wrong number, misunderstand a document, omit a relevant source or apply an inconsistent assumption. A more experienced reviewer could often identify the problem because it conflicted with other evidence or revealed a visible gap in the preparer’s reasoning.
Generative AI produces a different kind of error. It can generate a false source with a plausible author, credible title, appropriate journal and correctly formatted citation. It can invent a case study that fits the argument and describe it in polished professional language. It can combine elements from several genuine sources into one reference that looks authentic.
The result may look more finished, rather than less finished, than a junior employee’s draft. That creates a weakness in the review chain. A manager may read the narrative on the assumption that the preparer checked the references. A senior manager may assume that the manager tested the evidence. By the time the work reaches a partner, several people may have reviewed the document without anyone opening the underlying source.
A fabricated reference can therefore survive several layers of competent human review.
The model therefore requires new review protocols targeted specifically at AI-generated failure modes. A hierarchy designed to review human reasoning cannot safely assume that polished AI-assisted work carries the same visible indicators of error.
AI-assisted work can also produce what Wayne Banks described as “automation deference”: a gradual loss of professional scepticism when a system repeatedly generates plausible, well-formatted output.
Reviewers are likely to challenge an obviously rough draft. A polished report can feel almost complete, encouraging them to concentrate on its conclusion, commercial message or presentation rather than testing whether each source exists and supports the claim being made.
What proper source verification requires
Telling employees to “check AI output” is not enough. For material work, the firm needs a record of who checked what.
At a minimum:
Each material factual claim, case study and citation should lead back to an identifiable original source.
The person responsible for checking that source should be named in the review record.
The reviewer should confirm both that the source exists and that it supports the claim made.
The file should show where AI materially influenced the research, evidence, analysis or drafting.
Partners and senior reviewers should be told which parts of the work were AI-assisted.
High-risk reports should undergo independent source verification before publication or delivery.
Firms should monitor failed checks and corrected errors, not only adoption and productivity.
Responsibility should remain with the firm and the responsible partner or publication owner; it should not be displaced onto the person who entered the prompt.
The amount of checking required will depend on the stakes.
A short internal summary does not require the same process as an assurance report affecting welfare recipients, a government healthcare strategy or advice that may influence regulated decisions. But a public report carrying a Big Four brand should at minimum be able to demonstrate that its core claims and references were checked.
The failures are no longer isolated
There is no evidence that Paul Griggs’s comments caused the errors in the PwC Middle East reports. Some of the reports pre-dated his March 2026 interview, and PwC Middle East and PwC US are separate member firms within the global network.
The relevance of his remarks is structural rather than causal.
They demonstrate the intensity of the pressure surrounding AI adoption at the highest level of the Big Four at the same time that repeated failures are exposing weaknesses in the systems used to verify AI-assisted work.
The firms may possess responsible-AI frameworks, internal policies and approved tools. But policies should be judged by what they prevent.
Across the four networks, fabricated or distorted material has now reached government departments, clients and the public. Reports have been corrected or withdrawn. Deloitte refunded part of a government fee. An accounting regulator is investigating the Canadian Deloitte report.
The recurrence across firms, countries and publication types makes it increasingly difficult to attribute these failures to a few careless employees.
The Big Four are introducing a technology capable of generating authoritative falsehoods into an operating model built on delegated preparation, hierarchical review and trust in the work performed at the level below.
At the same time, they are telling employees and partners that extensive AI adoption is essential to productivity, career progression and the firms’ future commercial relevance.
That combination requires controls more explicit and demanding than general instructions to exercise judgement and check the output.
PwC has completed the Big Four set. Unless the firms make verification as concrete as adoption, it is unlikely to be the last Big Four AI failure.
Trying to keep up with technology at Deloitte, PwC, EY and KPMG? Start here →
https://www.big4news.com/t/technology-and-ai
About Claudine Cassar
I’m a corporate anthropologist and former Deloitte equity partner. I sold my technology business to Deloitte in 2016 and led the Malta Consulting team for five years. I am the founder and editor of Big4News, which provides independent, clear analysis of PwC, Deloitte, EY, and KPMG — free from corporate spin.
Find me on LinkedIn, X, Instagram, or my author website.
Feel free to reply to this newsletter — I read every reply.








