Key Takeaways
AI is eroding the junior work that traditionally developed audit judgment.
Training must replace some of the learning once gained through repeated engagement work.
The Big Four pyramid is both a staffing model and a training model — AI challenges both.
Firms may need to rethink how they staff, promote and reward future auditors.
What Arthur Andersen understood about training
I spent part of my early career at Arthur Andersen & Co., writing what were formally called Foundation Skills Units — FSUs — though everyone just called them Green Books. They were mandatory self-study materials, built to be completed before staff ever set foot in Andersen’s training center in St. Charles, Illinois.
Depending on the role, onboarding could run for weeks, sometimes months. The Green Books covered the foundations: industry knowledge, project management, communications, the basic technical and consulting skills every new hire needed before they could be trusted on a live engagement.
The point of that system wasn’t rigor for its own sake. It was calibrated to a specific, practical outcome: staff had to leave that training able to perform on day one. Not eventually. Not after a supervised ramp-up period measured in months. Day one.
I think about that standard every time I read another firm’s account of how it’s redesigning training for the AI era — including KPMG’s recent overhaul of its audit internship program, which sent nearly a thousand interns through an intensive experience at its Lakehouse facility in Florida, built around fraud scenarios, ambiguity, and judgment under incomplete information.
It’s a serious, well-intentioned response to a real problem. But it invites a question Andersen’s system answers more cleanly than most of what’s replacing it: what, exactly, does this training produce, and how do we know?
I should say plainly what many readers of this publication already know: AA&Co ultimately became engulfed in serious audit failures and controversies. But those failures do not mean that every system the firm developed was without value. That distinction is worth holding onto here, because it’s easy to let one discredit the other. It shouldn’t, because the firm’s approach to training offers lessons that remain relevant as firms rethink how auditors learn.
Two different disciplines, one training problem
Most interns arrive at a firm with a genuinely strong accounting foundation. They know debits and credits, revenue recognition, the standards. What they don’t yet know is auditing itself, which is a different discipline entirely. Auditing means gathering evidence, exercising professional judgment, evaluating risk, and getting comfortable making decisions when the facts in front of you are incomplete or contradictory.
Even Andersen’s Green Books, for all their rigor, were mostly building the first layer — the accounting and consulting fundamentals, not the audit-specific judgment that sits on top of it.
That second layer has traditionally been taught a different way: not in a classroom or through self-study, but by doing the work. A1s and A2s performed routine testing, under supervision, over and over, on real engagements with real stakes. The repetition is what built the judgment. No classroom curriculum could fully replicate it because the engagement itself was part of the curriculum.
That’s precisely the layer AI is now eating into. And it’s worth being honest about what that means: the mechanism that used to help teach professional skepticism is being eroded at the same time the need for that skepticism is going up, not down.
What the data actually says
KPMG’s own research, a field study run jointly with the University of Texas at Austin, gives us a rare look at the scale of that challenge. In an experiment involving 523 early-career professionals working alongside AI agents, roughly half outperformed what the AI could produce alone. About a quarter produced results comparable to the AI-only baseline. And about a quarter performed worse than the AI working by itself.
Sit with that last number for a moment. A quarter of the participants, doing assignments meant to resemble real client work, produced results inferior to a machine working unsupervised.
Those results illustrate the difficulty of the problem facing audit firms. Giving junior professionals sophisticated AI tools does not automatically produce better judgment. The unresolved question is whether intensive training and simulation can develop, at scale, capabilities that were previously built through years of supervised engagement experience.
At Andersen, the intended output was explicit: staff were expected to arrive on an engagement prepared to perform, rather than beginning a months-long process of learning the fundamentals on the job. I haven’t yet seen firms publicly define an equally concrete performance standard for this new generation of AI-era training — or report whether participants actually meet it.
That’s the standard I’d hold this kind of program to, and it’s the standard the industry should be asking of itself before declaring the problem solved.
Tim Walsh, KPMG US's Chair and CEO, recently made essentially the same argument about the human skills that will matter. He pushed back on what he called the biggest misconception about AI — that it makes people less important — arguing the opposite is true: as routine work shifts to AI, the skills that create value become more distinctly human, among them judgment, adaptability, critical thinking, leadership, and the ability to build trusted relationships. That’s the reasoning behind, in his words, “investing in our interns differently.”
I don’t disagree with a word of that. But it’s worth noticing what the statement doesn’t address. It’s a claim about what training should produce. It says nothing about whether the firm’s staffing model, promotion criteria, and economics are prepared to reward that output once training produces it.
The economics nobody wants to say out loud
There’s a reason firms have been slow to confront this directly, and it isn’t a lack of good intentions. It’s the leverage model.
The Big Four’s economics have always run on a pyramid: a large base of relatively inexpensive A1s and A2s performing billable, hands-on work, supervised by a much thinner layer of managers and partners above them. That structure did two things simultaneously. It generated margin, because A1 and A2 hours were profitable to bill. And it trained the next generation of senior staff, because the work those junior associates performed was also how they learned to eventually review, challenge, and sign off on other people’s work.
AI is quietly breaking the first half of that arrangement while making the second half more urgent. If routine testing no longer needs a large A1/A2 cohort to execute it, the commercial case for stacking big classes of junior associates onto engagements weakens, even as the profession needs those same associates to develop sharper judgment, faster, than the old apprenticeship model ever asked of them. You can’t simply keep the pyramid’s shape and swap out what fills the bottom layer. The shape itself was built around a kind of work that is going away.
That has real consequences for firm culture, and it's the part of this conversation I'd like to see get more attention. Fewer A1 and A2 hires, if that's where this goes, means different partner-to-staff ratios, different promotion timelines, and a different definition of what makes someone promotable in the first place.
Professional-services firms have traditionally had relatively visible measures of contribution: utilisation, hours worked, engagements completed and increasing levels of responsibility. Judgment is harder to observe and harder still to measure consistently. If judgment becomes a larger part of the value junior professionals are expected to add, firms will need stronger ways of identifying, developing and rewarding it. That's a much harder measurement problem than anything training design alone can solve, and it's one the profession hasn't started answering publicly yet.
What the future auditor actually looks like
I don’t think the answer is more hours of training, or even better-designed simulations, though both are worth pursuing. I think the profession needs to be honest that it’s solving two problems at once — how to teach judgment without the repetition that used to teach it for free, and how to restructure a business model that was never designed to reward judgment as its primary output.
The future auditor won’t be defined by how much work they can perform. They’ll be defined by the quality of the questions they ask and the judgment they apply to the evidence AI puts in front of them. Firms that treat this as a training problem alone may find that even a well-designed program cannot, by itself, solve the deeper challenge of developing judgment as routine work moves to AI. The firms that also rethink how they staff, promote, and reward around judgment — rather than volume — are the ones actually answering the question this moment is asking.
Arthur Andersen & Co understood, decades ago, that training needed a clearly defined practical outcome. The industry owes itself the same clarity now.
This article is part of the Big4News Analysis and Technology & AI series.
Who is William Englehaupt?
William Englehaupt is an audit process-improvement specialist with more than 35 years of experience spanning professional services and organisational transformation. He held senior roles at PwC, KPMG and Accenture as well as earlier experience at Arthur Andersen.
At KPMG US, where he was a director from 2018 to 2023, Englehaupt worked in the firm’s audit quality improvement programme, leading workshops and improvement initiatives involving engagement teams, audit specialists and other KPMG member firms. Before that, he spent more than three years at PwC, where he worked on Assurance transformation, engagement quality and continuous improvement, including Lean and Six Sigma initiatives.
Earlier in his career, Englehaupt held senior roles at Accenture, George Group and KPMG/BearingPoint, and also worked with the United Nations as a Senior Programme Officer and international consultant.
Englehaupt is a Lean Six Sigma Master Black Belt. He is not a CPA and approaches audit principally from an operational and process-improvement perspective, with particular interests in audit planning, workflow, review processes, rework and project management.
He founded Auditology in 2025, where his work focuses on audit process improvement and the application of Lean Six Sigma and systems thinking to audit execution. He is the co-author, with Brian Bilsback, of Managing the Hell Out of Your Audit (2025), and the author of The Productive Auditor (2026).







