Tuesday 25 August,ย Workshop 1, 5:30 pm โ 6:30 pm. Jesus College, Cambridge University (these are my lecture notes for today’s session, matching my slide deck).
Professor William Byrnes, Texas A&M University School of Law; Co-author, Money Laundering, Asset Forfeiture and Recovery and Compliance: A Global Guide (Lexis); Co-author, FATCA & CRS Compliance (Lexis)
Howdy!
Introduction: The Paradox and its Counter Paradox
Tax authorities today face a paradox. We possess more information than any revenue administration in historyโelectronic returns, e-invoices, customs declarations, financial reports, property records, beneficial-ownership data, automatic exchanges of information, and digital-asset reporting. Yet the volume and complexity of that information exceed the practical capacity of human auditors to analyze it unaided.
Artificial intelligence addresses that mismatch. It can identify patterns across millions of records, prioritize cases, trace relationships among entities, detect transactions inconsistent with economic reality, and help auditors retrieve and summarize evidence. AI is therefore not merely an automation tool. It is a mechanism for converting data into enforcement capacity. Based on the research and analysis of my Texas A&M colleague, Pramod Kumar, and me, we identify three principal drivers:ย (1) efficiency gains, (2) resource optimization, and (3) improved detection of fraud and evasion.
But the same capability creates a second paradox. The more effective a tax authority becomes at observing, profiling, and predicting taxpayer conduct, the greater its responsibility to explain how state power is being exercised. A tax authority can be technically accurate and still act unlawfully. It can increase audit yield while producing discriminatory outcomes. It can place a human at the end of an automated process while giving that person no realistic ability to challenge the machine.
So the question for this class is not, โShould tax authorities use AI?โ The evidence shows that they already do. The better question is:
How can a tax authority use AI to improve compliance while preserving legality, taxpayer rights, human judgment, and public trust?
The OECD describes tax authorities as experienced public-sector users of AI, particularly in fraud detection, risk assessment, decision support, and taxpayer service. The IMFโs guidance for senior officials similarly treats AI adoption as an operational and governance decision requiring legal and ethical analysis, use-case assessment, strategy, and risk management.
Part I โ What โAIโ means in tax administration
For todayโs purposes, AI is a machine-based system that infers from its inputs how to generate predictions, recommendations, content, or decisions. In tax administration, that definition embraces a family of tools rather than a single technology: rules engines, supervised and unsupervised machine learning, natural-language processing, network analysis, computer vision, generative AI, and emerging agentic systems.
IA. Five Levels of Analytical Capability
It is useful to distinguish five levels of analytical capability:
- Rules-based automationย applies an explicit test: for example, flag a refund above a threshold where required third-party documentation is missing.
- Supervised learningย studies previously labeled audits or fraud cases and predicts which new cases resemble productive historical cases.
- Unsupervised learningย searches unlabeled data for clusters, outliers, and relationships that officials did not specify in advance.
- Generative AIย summarizes, retrieves, translates, or drafts material from a controlled knowledge base.
- Agentic AIย can plan and execute multi-step tasks with greater autonomyโfor example, gathering records, reconciling transactions, preparing an issue list, and routing the case for review. Greater autonomy must mean stronger authorization limits, logging, and intervention controlsโnot less oversight.
IB. The Distinct Tax Audit Steps: Selection, Investigation, and Decision
A fundamental distinction is amongย selection,ย investigation, andย decision. A model may select a return for review. It may then help an auditor gather or organize evidence. But the ultimate assessment, penalty, collection measure, or referral for prosecution remains a legal decision. We should resist designing one undifferentiated โAI enforcement system.โ Each stage has different evidentiary standards, legal authority, and potential consequences.
With that vocabulary in place, let us follow an AI system from raw data to an enforcement outcome.
The following diagram illustrates the essential separation between lawful inputs, machine analysis, meaningful human judgment, proportional response, taxpayer recourse, and model monitoring.

Figure 1: A responsible tax-AI lifecycle separates model-generated risk signals from human decisions, taxpayer safeguards, and continuous independent monitoring.
Part II โ The practical AI toolkit
1. Risk scoring and predictive audit selection
Eight years ago, my colleague here today, Dr. Dionysis Demetis, and I gave a similar workshop: tax auditing with big data and machine learning. Our main point in that workshop was to show the capabilities of big data analytics, as they existed in 2018, for risk scoring to allocate government resources.
Risk scoring asks: Which returns, transactions, or taxpayers merit scarce human attention? Inputs may include filing history, sector benchmarks, related-party dealings, customs activity, e-invoices, third-party reports, prior examinations, and payment behavior.
The output should not be โthis taxpayer is guilty.โ It should be a ranked lead accompanied by reason codes, confidence information, and the relevant discrepancies. A defensible model might say: reported gross margin materially differs from comparable filers; purchases reported by counterparties exceed purchases claimed; and declared payroll is inconsistent with operational scale. That is actionable intelligence. A bare score of โ94โ is not.
Case Use Example: The U.S. IRS supplies a useful current example. TIGTA reports that the IRS uses AI-related models for individual-return classification, corporate line-anomaly analysis, and large-partnership selection. The Line Anomaly Recommender estimates expected relationships among corporate-return line items and assigns risk based on deviations; the Large Partnership Compliance model identifies outliers across complex partnership filings. Importantly, outputs undergo human classification and expert review.
The lesson is not that AI should maximize assessments. It should reduceย unproductive audits and unnecessary taxpayer burdenย while preserving representative or random audit programs needed to understand the compliance population. TIGTA found that the IRS had not yet established sufficient processes to demonstrate that some AI models outperformed prior methods in real-world conditions. It recommended incorporating feedback from examination outcomes, defining performance metrics, using ensemble methods where appropriate, and monitoring for model drift.
2. Anomaly detection and cross-database matching
Anomaly detection asks a different question: What does not fit? An unsupervised model may compare a taxpayer with its own history, its industry, or economically similar entities. It may detect:
- revenue growth without corresponding labor, inventory, or consumption changes;
- deductions that move independently of business activity;
- VAT purchases without corresponding seller declarations;
- a refund claim inconsistent with the taxpayerโs supply chain;
- property acquisitions inconsistent with reported income; or
- repeated changes in identity, address, director, or bank account associated with short-lived companies.
Case Use Examples: Spainโs approach illustrates cross-database matching: property ownership, land-registry information, estate-agent records, and online rental advertisements can be compared with tax filings. The UKโs data environment similarly combines property, travel, border, vehicle, and rental information to identify undeclared income or unexplained asset acquisition.
BUT caution is critical! This, Pramod Kumar and I argue, is a critical element of good AI tax audit governance:ย an inconsistency may reflect timing, classification, erroneous third-party data, or a lawful business explanation.ย
CAUTION: The AI system should preserve the underlying source and permit correction rather than convert mismatch into presumption. [Minority Report movie sci-fi reference. But sci-fi turned real life: UK Postmaster scandal that, to this day, accountability has not been upheld]
3. Network and graph analysis
Many sophisticated schemes are not visible at the taxpayer level. They exist in the relationships among taxpayers.
Graph analysis represents taxpayers, companies, directors, addresses, bank accounts, invoices, customs declarations, and beneficial owners as nodes connected by transactions or affiliations. It can reveal circular invoice chains, repeated use of common contact information, rapidly changing shell entities, and clusters that move credits or losses without a plausible commercial pattern.
Case Use Example: This is especially valuable for VAT carousel fraud and false invoicing. A traditional audit might examine one invoice. But my colleague today Dr, Demetis was one of the very first experts in Graph and Relationship analysis which asks: (1) whether the invoice sits inside a circular or rapidly expanding network; (2) whether the supplier lacks employees, premises, or corresponding purchases; (3) whether the same bank account or address appears across nominally independent firms; and (4) whether value repeatedly returns to its origin.
If you saw our workshop in 2018 wherein Dr. Demetis presented his heat mapping exercise based on a substantial pool of the bank accounts for a EU member state, and movements among accounts, youโll appreciate the mass data analytics capability of machine learning, but that it must be translated into something โusableโ by our human analytical abilities, such as visual representation via a graph and relationship analysis.
Case Use Example: Chinaโs Golden Tax system uses data mining and anomaly detection across mandated e-invoice data to identify suspicious VAT refunds, shell companies, and illicit fapiao invoicing networks. Brazilโs electronic invoice and SPED bookkeeping environment similarly enables large-scale anomaly analysis, with human auditors validating findings before notices are issued.
4. Computer vision, images, and geospatial analysis
AI can also analyze what is visible but not declared.
Case Use Study: Greeceโs Independent Authority for Public Revenue has used satellite-image analysis, that we spoke about in our 2018 workshop, to identify potentially undeclared swimming pools and property improvements and compare them with declarations. Image recognition can similarly support property-tax mapping, construction monitoring, agricultural assessments, and customs inspection.
CAUTION: Pramod Kumar and I caution in our forthcoming book: The legal and operational controls should be explicit. What imagery may be used? Is it public, acquired under statute, or purchased from a vendor? How recent and accurate is it? Can the model distinguish permanent improvements from temporary objects or visual artifacts? Before assessment, an official should verify the image, ownership, date, and governing tax rule.
5. Generative AI for auditors and taxpayers
Generative AI is most defensible where it reduces cognitive burden without determining liability. A secure system can summarize contracts, extract clauses, compare a taxpayer submission against an information request, retrieve relevant guidance, draft a factual chronology, translate correspondence, or create a first draft of an examination memorandum.
For taxpayer service, virtual assistants can answer routine questions, route correspondence, explain filing obligations, and identify omissions before submission. These functions can improve compliance because many errors arise from complexity rather than intent. Your materials emphasize chatbots and workflow support alongside enforcement; a reminder that AI should make lawful compliance easier, not merely make enforcement more powerful.
But a generative answer must not silently become official law. The system should use approved sources, display citations, distinguish authoritative law from guidance, preserve prompts and outputs where appropriate, protect return information, and require human approval for consequential communications.
Part III โ How AI exposes particular forms of tax evasion
- VAT and refund fraud
AI can reconcile invoice-level seller and purchaser records; detect unusual refund ratios; trace circular trading; identify missing traders; and rank networks by expected fiscal exposure. China and Brazil demonstrate the value of e-invoice infrastructure, while India has applied transaction-network analysis to shell or โbogusโ firms issuing false invoices.
The practical sequence should be: mismatch โ network analysis โ evidence retrieval โ human verification โ proportionate intervention. A prompt asking a taxpayer to correct a discrepancy may be appropriate before a full audit where the risk and likely harm are limited.
- Offshore noncompliance
Automatic exchange of information makes foreign account data available; AI can then match names and entities, identify non-filers, reconcile balances with declared income, and prioritize cases. Based on my research for my Lexis tax treatise called FATCA & CRS Compliance, I reported here in my 2024 Cambridge lecture about the impact from 123 million financial accounts and EUR 12 trillion of assets shared within the referenced automatic exchange of information environment; and I described as a case study about Peru receiving information from 40 jurisdictions concerning 43,000 citizens and 57,000 foreign accounts.
CAUTION: Yet ownership matching is probabilistic. Transliteration, joint ownership, trusts, and duplicated names can produce false links. Entity-resolution confidence must therefore be reviewable before contact or assessment.
- Transfer pricing and complex groups
AI can compare related-party margins, customs values, country-by-country patterns, royalty flows, service charges, and year-to-year changes. Network tools can map legal ownership against functional and transaction flows. But officials should make the classroom distinction explicit:
A transfer-pricing anomaly is not itself evasionโand tax avoidance is not automatically fraud.
AI identifies questions: persistent losses in a distributor, margins outside a defensible range, unexplained payments to low-tax affiliates, or trade values inconsistent with comparable goods. But it should not be relied upon for โjudgmentโ. Human auditors must still apply the law, perform functional analysis, evaluate comparability, and distinguish error, dispute, avoidance, and intentional evasion. The source materials emphasize the scale and complexity of intercompany transfers and the value of specialized audit capacity rather than treating every adjustment as proof of misconduct.
Part IV โ Comparative governance models
| Jurisdiction | Operational approach | Principal governance lesson |
| European Union | GDPR, administrative law, and the EU AI Actโs risk-based framework operate together. | GDPR Article 22 protects against solely automated legally significant decisions and provides safeguards including human intervention and contestability. Tax enforcement is not expressly listed in Annex III of the AI Act, so classification requires careful use-case analysis; nevertheless, its controlsโdata quality, logging, documentation, human oversight, accuracy, and cybersecurityโare a useful benchmark. |
| United States | AI supports classification and audit selection within a broader framework of return confidentiality, the Privacy Act, administrative law, appeals, and judicial review. | Oversight is predominantly institutional and ex post. TIGTA stresses outcome measurement and drift monitoring; GAO reported 126 active IRS AI use cases as of June 2025 but identified incomplete inventories, skills gaps, and lack of agency-wide investment management. |
| United Kingdom | HMRC Connect aggregates diverse data and generates intelligence for human officers; the jurisdiction favors sectoral and transparency-oriented governance. | Consequential outputs must remain explainable and contestable; publication of algorithmic transparency records offers a replicable accountability mechanism. |
| India | AI/ML-supported scrutiny selection draws on income, GST, high-value transaction, property, and securities data. | India relies more heavily on constitutional privacy, equality, natural justice, notification, assessment participation, and judicial review than on a GDPR-style statutory right against automated decisions. |
| China | Golden Tax uses e-invoice analytics, risk alerts, and cross-checking with customs, bank, and VAT data. | PIPL Article 24 emphasizes transparency, fairness, impartiality, explanation, and protections against unreasonable differential treatment, though oversight is more centralized and administrative. |
| Brazil | Receita Federal uses electronic invoice and SPED analytics while auditors validate findings. | LGPD Article 20 supplies review and explanation rights for decisions based solely on automated processing; the model combines advanced data use with contestability. |
| Japan | Data analytics and taxpayer-assistance tools operate under privacy law, administrative review, and largely soft-law AI guidance. | Executive responsibility, professional norms, and human validation can support accountability, but soft law should not substitute for clear remedies where consequences grow. |
Two Dutch cases should shape every implementation discussion. In SyRI, a court rejected an opaque welfare-fraud risk system on privacy and human-rights grounds. In the childcare-benefits scandal, algorithmic profilingโincluding problematic use of nationalityโcontributed to wrongful treatment of thousands of families and prolonged difficulty obtaining human correction. These were not merely defective models; they were defective administrative systems in which suspicion hardened into adverse action and review failed.
I already mentioned the UK Postmaster scandal. The UK Postmaster scandal involved accounting software developed by Fujitsu corporation that turned out to be โbuggyโ to say the least. In briefest summary, the system generated a consistent flow of false positives that local office postmasters were stealing. The UK government prosecuted over 900 postmasters for theft and fraud over a six-year period based on the software outputs, bankrupting most of these postmasters, imprisoning many; at least 13 postmaster suicides were attributed to these prosecutions. We only know now that the system was, as I politely call it, buggy, because of intense investigative journalism. ย ย ย ย
U.S. academic research has focused on bias evidence, raising a related concern: an algorithm need not use race explicitly to create racial disparity. A comparative survey reports audit-rate disparities associated with models that emphasize certain low-income credits. The governance response is not to evaluate intent alone; it is to measure false-positive rates, audit burdens, no-change outcomes, and other effects across legally appropriate taxpayer segments, and only then investigate causal pathways.
Part V โ A practical implementation blueprint
So let me bring my address to a close, and then we can dive into the workshop discussion. Pramod Kumar and I recommend ten controls for any tax authority deploying consequential AI.
- Establish statutory and policy authority.ย Document the legal basis, purpose, permissible data, decision stage, retention period, and responsible official.
- Maintain an enterprise AI inventory.ย Include internally developed, vendor-provided, sensitive, generative, and experimental systems. GAOโs IRS review demonstrates that incomplete inventory information prevents strategic oversight.gdpr-info+1
- Conduct an algorithmic impact assessment.ย Evaluate privacy, equality, due process, security, accuracy, affected populations, alternatives, and remedies before deployment.
- Use representative and legally permissible data.ย Document provenance, missingness, label quality, and proxy variables. More data are not always better data.
- Define success before piloting.ย Measure additional productive leads, no-change audits, false positives, processing time, taxpayer burden, appeal reversals, and distributional effectsโnot gross assessments alone.
- Require meaningful human review.ย The reviewer must understand the reasons, access underlying evidence, possess authority to override, and record the independent decision.
- Provide explanation and correction.ย Protect sensitive anti-evasion parameters, but communicate the substantive discrepancy, data sources that may be challenged, and route to reconsideration.
- Create model documentation and logs.ย Maintain model cards, versions, training periods, features, thresholds, reason codes, overrides, access records, vendor changes, and incident reports.
- Monitor outcomes and drift.ย Compare model recommendations with completed audit results; test error rates and burden; retrain or suspend models when laws, taxpayer behavior, or data relationships change. TIGTA made precisely this feedback-and-measurement recommendation.tigta+1
- Preserve institutional capability.ย Pair auditors, lawyers, data scientists, privacy officers, cybersecurity professionals, taxpayer-service experts, and appeals personnel. Procurement must not transfer practical control of public authority to an opaque vendor.
Inevitably, the Ombudsman Office of a Tax Authority must have the authority and resources necessary to perform its role in the context of AI leveraging for audit. In the USA, we have the Taxpayer Advocateโs Office and TIGTA. But both would need a budget to be able to employ computer engineering technicians capable of โkicking the tiresโ and challenging AI systems to determine annually whether it is doing what we want it to do, not doing what we do not want it to do, and whether appropriate governance and legal safeguards, like taxpayer rights, are being exercised. ย ย
QUESTION: Who, by name or office, is accountable when the model is wrongโthe model owner, data steward, auditor, vendor, chief information officer, or agency head? If the organization cannot answer that question before deployment, it is not ready to deploy.
Part VI โ The next frontier
[Dionysis – this is your bailiwick so feel free to jump in] Generative and agentic systems will move beyond scoring into multi-stage case preparation. They may collect approved data, reconcile documents, identify issues, draft correspondence, schedule follow-up tasks, and learn from outcomes. That could greatly increase auditor capacityโbut it also creates risks of hallucinated facts, unauthorized tool use, over-collection, automation bias, and actions taken before an official notices.
The appropriate architecture isย bounded autonomy: approved data only; role-based tool permissions; no autonomous assessment, penalty, seizure, disclosure, or referral; mandatory citations to source records; complete logs; confidence and exception reporting; and a readily available stop mechanism. The IMF, for example, treats AI strategy, use-case risk assessment, policy, and operational introduction as inseparable.
International coordination is also advancing. The U.N. Subcommittee on Tax Administration and Artificial Intelligence, established in October 2025, is preparing a practical guide covering fraud detection, risk assessment, taxpayer service, governance, data and gender bias, integrity, and confidentiality, with a draft due by October 2027 and a particular focus on developing countries. The OECD is examining AI across tax administration within its broader work on trustworthy government AI, while the IMF provides a use-case risk-assessment methodology for senior officials.
Finally, in our article Taxing the Intelligence Age published by Tax Notes, my colleague Prof. Pramod Kumar and I argued for a wider policy warning: AI policy can become fragmented when different levels of government pursue overlapping instruments without a common base, allocation rule, or coordinating mechanism. The administrative analog is clearโdo not allow every audit division to procure its own model, use incompatible definitions, and create disconnected accountability regimes.
Conclusion
Artificial intelligence can help tax authorities see what was previously invisible: a hidden invoice network, an offshore account mismatch, an unexplained property improvement, an anomalous partnership structure, or a pattern distributed across millions of transactions.
But AI does not determine what the law means. It does not decide whether evidence is sufficient. It does not supply proportionality, judgment, compassion, or procedural fairness.
The proper formula is therefore:
High-quality data, narrow lawful purpose, explainable analytics, meaningful human judgment, proportionate intervention, taxpayer recourse, and continuous accountability.
The best AI-enabled tax authority will not be the one that generates the highest number of audit flags. It will be the one that identifies genuine noncompliance more accurately, reduces unnecessary burdens on compliant taxpayers, explains its actions, corrects its mistakes, and earns public confidence.
Our governance takeaway for today: AI should strengthen the rule of law, not become an alternative to it.
Discussion questions for tax officials in workshop today
- Which compliance problem in your jurisdiction has sufficiently reliable data to justify an AI pilot?
- What should count as success: additional assessments, collections, fewer no-change audits, deterrence, or reduced taxpayer burden?
- At what point must a taxpayer be told that AI materially influenced the case?
- What explanation can be provided without disclosing parameters that would enable gaming?
- How will the authority test for indirect or proxy discrimination?
- Which actions must never be delegated to an autonomous system?
- Can the designated human reviewer genuinely reject the modelโor will the reviewer merely rubber-stamp it?
- How should closed audits, appeals, and reversals feed back into model validation?
- What capability must remain in-house when a system is vendor-supplied?
- If an AI-supported action is wrong, who is accountable and what remedy will the taxpayer receive?









Robert Bloink, J.D., LL.M.


condition for a finding of State aid, that the Netherlands-Starbucks APA conferred a selective advantage on Starbucks’ Netherlands manufacturing subsidiary (SMBV, aka the โroasting operationโ) that resulted in a lowering of SMBVโs tax liability in the Netherlands as compared with what SMBV would have paid under the Netherlandsโ general corporate income tax system dealing with third parties.
