How AI in Tax Law and LLMs Can Help Us Better Understand U.S. Tax Code
Episode 43 • 15th December 2025 • Beneath the Cypress and Star • BlueRidge Pundit
00:00:00 00:12:03

Share Episode

Shownotes

Summary

AI in tax law offers a way to make dense public information easier to explore. This episode examines how large language models could support research, compare provisions, and explain the U.S. tax code while keeping legal interpretation and policy choices accountable to people. It connects access to information with economic inequality and the distribution of resources. The opportunity is substantial, but reliable research requires more than a fluent answer.

Key Takeaway

  • AI for tax research can assist with summarizing, comparing documents, and locating material for review.
  • Statutes, regulations, court decisions, and agency guidance have different roles and must be read in context.
  • BillSum is a legislative summarization dataset, while compliance analytics and legal research tools perform different tasks.
  • A useful explanation identifies its sources, relevant dates, uncertainties, and questions that still require expert judgment.

How Can AI Explain Tax Law?

A reader can use an LLM to turn a passage into a plain-language explanation, identify defined terms, or compare two supplied versions of a provision. Those tasks can make an unfamiliar document more approachable. They do not establish whether the resulting interpretation applies to a particular taxpayer or whether an omitted exception changes the answer.

The Internal Revenue Service’s guide to legal authorities distinguishes the Internal Revenue Code in Title 26 from Treasury regulations and other official guidance. It also stresses checking the tax year and effective dates. Research needs that context: a correct quotation from a newer provision can still be the wrong authority for an earlier year.

LLMs and Tax Code Research

The episode explores the possibility of connecting legislation with its history and consequences. A VLDB workshop paper by Andrea Colombo and Francesco Cambria describes an LLM-assisted process for organizing U.S. legislation and links between provisions. Mapping citations can help researchers navigate related material. It does not, by itself, determine which policy is fair or prove the economic effects of an amendment.

Another example requires a clear distinction. The Association for Computational Linguistics paper introducing BillSum describes a dataset of U.S. Congressional and California bills for summarization research. It is not a comprehensive tax-advice platform. Comparing these projects by their actual purpose makes their contributions easier to assess.

Legal AI and the Limits of Automation

A Stanford Law School study of AI legal research tools found that retrieval-based systems could still generate incorrect information. The practical lesson is to open the cited authority, check whether it supports the claim, and separate the original text from the model’s interpretation. Confidence in the wording is not a substitute for evidence.

Enforcement analytics is a separate application. The IRS’s 2024 operating-plan update described using AI and advanced analytics to help select complex partnerships for examination. Selecting cases for human review is different from establishing liability, rewriting legislation, or showing that an LLM understands the entire system.

Transparency and Public Accountability

Better access to information could help people ask more informed questions about revenue, spending, and distribution. Our discussion of participatory budgeting and public decisions examines a complementary approach to participation. Capitalism and the distribution of value and rethinking progress through a well-being economy explore the values behind economic choices.

A sound research process preserves the documents being analyzed, identifies assumptions, and makes disagreements visible. AI can help organize a debate; elected institutions, legal professionals, and the public still have to evaluate its conclusions.

Frequently Asked Questions

Q1: How can AI explain tax law?

It can summarize supplied provisions, explain terminology, and suggest questions for further research. Check each explanation against the relevant legal sources and dates.

Q2: Can an LLM replace professional tax advice?

A general explanation cannot establish how the law applies to someone’s full circumstances. Individual decisions require current authority and qualified review.

Q3: Is BillSum a tax compliance product?

No. It is a research dataset for legislative summarization, with bills and reference summaries used to evaluate methods.

Q4: Can legal AI produce incorrect citations?

Yes. Research has found errors even in systems that retrieve legal documents. Verify citations and the propositions attributed to them.

Q5: Does organizing the tax code automatically make it fairer?

No. Organization can support analysis, but fairness depends on policy objectives, distributional evidence, public scrutiny, and accountable decisions.

Related Episodes

Sources & Further Reading

Transcripts

1

::

Today we are diving deep into the US legislative and tax code.

2

::

The ultimate labyrinth.

3

::

It really is a system that is famously convoluted and definitely stressful.

4

::

And let's be honest, a primary source of political fog.

5

::

It is.

6

::

And for centuries, lawmakers have had to approach any kind of reform in really the

7

::

only way they could piecemeal one piece at a time.

8

::

Exactly.

9

::

Just think about it.

10

::

No single person or office could ever review the entire legal framework.

11

::

Holistically, it's just too big.

12

::

So every new law was added into this, this fog you're talking

13

::

a dense historical fog.

14

::

And that often created these hidden, you know, unintentional consequences,

15

::

somewhere else on the code.

16

::

A fog is the perfect word for it.

17

::

And our mission in this deep dive is to show how large language models,

18

::

LLMs are moving way beyond being simple chatbots.

19

::

Much further.

20

::

They're becoming, in a way, forensic auditors for the law itself.

21

::

The promise here is clarity, pulling that system out of the fog and giving you the

22

::

listener a full panoramic view of how the system actually functions, not just how

23

::

We think it does; that holistic view must be what changes everything.

24

::

It's the key.

25

::

I mean, the central challenge from a technical standpoint is that the law

26

::

exists as just massive amounts of unstructured text.

27

::

We're talking about statutes, court rulings, hundreds of thousands of

28

::

pages of statutory text, decades of court rulings that sometimes even

29

::

conflict with each other, and then millions of pages of IRS guidance.

30

::

It's a mountain of data.

31

::

And LLMs are the first tool that can not just read that mountain of data, but

32

::

actually correlate it, see the connections.

33

::

Yes.

34

::

Correlated to system-wide scale.

35

::

Historically, if you wanted to know the impact of one little amendment, you'd

36

::

hire a team of lawyers.

37

::

And it's been months, maybe years, tracing citations.

38

::

Right.

39

::

Slow, expensive, and always incomplete.

40

::

AI transforms that whole process from this, you know, bureaucratic obstacle

41

::

into a data-driven diagnostic tool.

42

::

Okay.

43

::

So it can take a proposed policy change and check it against everything else at

44

::

once simultaneously.

45

::

It checks it against the statutory text, the relevant court cases, and even the

46

::

economic impact data.

47

::

And that must be the aha moment for lawmakers.

48

::

Suddenly the AI is flagging things.

49

::

Exactly.

50

::

It surfaces contradictions, redundancies, loopholes, all the outdated provisions

51

::

that human review just misses because the scope is too wide.

52

::

You can finally see what works and what doesn't.

53

::

And that's the real focus here.

54

::

The goal isn't just automation.

55

::

It's about redesigning the system to be fair.

56

::

By looking at all this data and case law outcomes, you can ground reform in evidence.

57

::

So it's about getting away from political guesswork.

58

::

And towards a system that actually reflects transparency, coherence and, you

59

::

know, equity at scale, because it's finally based on verifiable data.

60

::

Okay.

61

::

Let's get under the hood.

62

::

That scale of analysis sounds almost impossible.

63

::

How do LLMs take this messy pile of documents and turn it into something

64

::

structured?

65

::

The answer is something called a knowledge graph or a KG.

66

::

The KG is the scaffold, then.

67

::

It is the essential scaffold.

68

::

But to build it, to build the US legislative graph, researchers first hit this,

69

::

This huge and kind of surprising challenge was document quality.

70

::

The source documents for US laws are all over the place, format-wise, depending

71

::

on when they were published.

72

::

Let me guess, the older the document, the worse the quality.

73

::

You got it.

74

::

The oldest chunk from 1951 to 1992 is about 11,425 laws that only exist as image-based PDFs.

75

::

Oh, wow.

76

::

So just scans of old paper, ancient scanned images, which immediately brings up a

77

::

huge problem with optical character recognition.

78

::

Right.

79

::

OCR.

80

::

That's the software that tries to turn an image of text into actual text.

81

::

You can edit.

82

::

And if the scans are bad, the OCR is going to be full of errors.

83

::

We're talking typos, words split in half, punctuation just gone.

84

::

So your foundation is shaky from the start.

85

::

Very shaky.

86

::

Yeah.

87

::

So to solve this, they came up with an LLM-assisted strategy.

88

::

They basically turned the AI into a highly specialized editor.

89

::

And they used a specific model for this.

90

::

They used a powerful open-source model.

91

::

Llama 370B was instructed to act as that correction agent.

92

::

It reads the messy OCR text and just cleans it up.

93

::

It fixes the typos and keeps the structure.

94

::

All while trying his best not to hallucinate or make things up, that one step

95

::

turns that low-quality junk into high-quality text.

96

::

That's your node in the graph.

97

::

Okay.

98

::

So now we have clean text, clean nodes.

99

::

Now you have to connect them all.

100

::

That's the next step: edge extraction, finding all the relationships between those

101

::

thousands of laws, which must be hard because they don't just use clean hyperlinks, do they?

102

::

Not at all.

103

::

The references are often buried in unstructured text, like the act of 1974 or some short

104

::

title.

105

::

It's not consistent.

106

::

So you need an AI that can understand the legal intent of a sentence, not just match keywords.

108

::

Precisely for that, they fine-tuned the different smaller models, Mr.

109

::

All seven B, just for this one job: relationship identification.

110

::

And it classifies the type of relationship, the edge, you could call it.

111

::

That's right.

112

::

It's not just law A mentions law B.

113

::

It's how it mentions it.

114

::

So what are the categories?

115

::

There are four main ones.

116

::

Does one act amend another?

117

::

Does it abrogate it?

118

::

Which means repeal it.

119

::

Okay.

120

::

Is it the legal basis of another law?

121

::

Or does it just cite it for reference?

122

::

And running this process over 60 years of history.

123

::

That's what builds the graph.

124

::

That's what built it.

125

::

Yeah.

126

::

A massive, queryable graph with over 31,000 unique functional relationships: the edges.

127

::

That's incredible.

128

::

But what happens if the AI gets it wrong?

129

::

If Mistral misclassifies an edge.

130

::

That's a huge risk.

131

::

I mean, if it says a law merely cites another, but it actually

132

::

obligates it; that's a catastrophic error for a lawyer.

133

::

It would be.

134

::

And that's why these systems have to be transparent.

135

::

A misclassified edge could lead to an indefensible legal position.

136

::

The beauty of the graph is that a human can then query that edge

137

::

and check the original statutory language.

138

::

So the AI is the discovery tool, not the final word.

139

::

Exactly.

140

::

It finds the connection in seconds.

141

::

Then the human verifies it.

142

::

Got it.

143

::

Okay.

144

::

Let's shift from the theory to the real world.

145

::

How does this impact big finance?

146

::

Let's talk big tech and global taxes.

147

::

Right.

148

::

This is where that structured data gets applied to these incredibly dense financial documents.

149

::

I saw an AI legal agent was used to analyze three massive 10K filings,

150

::

Alphabet, Microsoft, and Nvidia.

151

::

It was the goal was to check their tax positions against the OECD's

152

::

Pillar 2 initiative.

153

::

That's the global minimum tax, the 15% threshold for the effective tax rate or ETR.

154

::

And these 10 key filings are hundreds of pages long.

155

::

The AI agent using this graph-based knowledge was able to run these really complex analyses.

156

::

And the results were specific.

157

::

I mean, very specific.

158

::

Incredibly, for example, it found that Nvidia's effective tax rate was

159

::

calculated to be 12%.

160

::

12%.

161

::

So that's three points below the global minimum.

162

::

Correct.

163

::

And that potentially exposes Nvidia to what are called top-up taxes.

164

::

They have to pay the difference that 3% to the jurisdictions where they operate.

165

::

And the AI to point to the exact page in the document to prove it.

166

::

With precise page citations.

167

::

That's the verifiability we're talking about.

168

::

The analysis also found these really different tax splits, right?

169

::

Domestic versus foreign.

170

::

It did.

171

::

Alphabet and Nvidia were heavily US-centric.

172

::

Over 90% of their tax burden was domestic.

173

::

But Microsoft was much more global, more like a 60 40 split.

174

::

That kind of insight used to take teams of people weeks to figure out.

175

::

And now it's minutes.

176

::

It's a huge shift, which brings us back down to Earth.

177

::

This isn't just for huge corporations.

178

::

It's becoming well essential for smaller law and accounting firms too.

179

::

Absolutely essential.

180

::

I mean, think about the pace of change.

181

::

Small firms get buried under this constant stream of new regulations,

182

::

interpretations and court decisions.

183

::

It just can't keep up.

184

::

They don't have the research staff.

185

::

But AI acts as an equalizer.

186

::

The efficiency gains are just massive.

187

::

How massive are we talking?

188

::

We're seeing an 80% reduction in time spent finding relevant tax guidance.

189

::

80%.

190

::

That's not just an improvement.

191

::

That's a total game changer for a small practice.

192

::

It really is.

193

::

That's time they get back to actually advising clients to use their professional

194

::

judgment.

195

::

It's a huge business advantage and better for the clients who get faster,

196

::

more accurate advice.

197

::

And now we're even seeing that clarity filtered down to the public with tools

198

::

like taxis.ai.

199

::

Right.

200

::

These are the user-friendly assistance programs that help individual taxpayers.

201

::

Yeah.

202

::

They use a technique called retrieval-augmented generation, or RAG.

203

::

A.A.G.

204

::

They draw on this knowledge base of up-to-date forms and instructions to give

205

::

you personalized guidance.

206

::

So it's demystifying a process that, frankly, a lot of people feel is

207

::

intentionally complicated.

208

::

It's a huge step toward equity, but this is really important, but we have to add a major caveat here.

210

::

Okay.

211

::

This is a powerful tool.

212

::

And like any powerful tool, it needs human judgment and validation.

213

::

It's an assistant, not a replacement.

214

::

So we have the good use case where agencies like the IRS use it for fraud

215

::

detection or bill analysis.

216

::

Right.

217

::

That's AI enhancing the existing process.

218

::

But then there's the other side, the push for pure efficiency, which brings us

219

::

to the, let's say, controversial use of the DOGE AI deregulation decision tool.

220

::

The name alone is something and its goal is huge, isn't it?

221

::

It's staggering.

222

::

The tool is supposed to analyze 200,000 federal regulations to identify and

223

::

eliminate a hundred thousand of them, a hundred thousand and the projected

224

::

savings are 3.3 trillion dollars annually, a very powerful incentive to move fast.

225

::

But what went wrong when they tested it?

226

::

You mentioned a HUD employee's experience.

227

::

Well, the employee who tested it reported that the AI's analysis just came to

228

::

the wrong legal conclusions.

229

::

How so it would sometimes misread the language.

230

::

It would claim that a regulation was outside the scope of the underlying

231

::

statute when, in fact, the AI had just misinterpreted the law I was based on.

232

::

So it's like the AI decided a key paragraph was irrelevant when that

233

::

paragraph was the entire legal foundation for the rule.

234

::

Precisely, which brings us back to this idea of a legal grade standard.

235

::

The best tools, the ones you can trust, are transparent and traceable.

236

::

You have to be able to follow the data trail back to the source text.

237

::

It has to be an enhancement to a professional's judgment, not a replacement for it.

238

::

What an incredible arc we followed from messy scanned PDFs to a precise,

239

::

queryable map of the entire US legal system.

240

::

LMS aren't just simplifying taxes.

241

::

They are making the underlying structure of the law visible for the very first time.

242

::

And that structural view gives us these profound insights.

243

::

Researchers analyzed the graph to see which subjects have the most influence.

244

::

They're centrality.

245

::

And what did they find about modern lawmaking?

246

::

They found that the centrality ranking of major subjects, things like economics, public finance, it's remained almost entirely unchanged since 2015.

248

::

Unchanged for nearly a decade.

249

::

Almost static.

250

::

Might they?

251

::

Which suggests a stabilization or maybe more provocatively, an

252

::

ossification of thematic focus in US lawmaking.

253

::

So the core plumbing of the system is stuck on the same things, even as the

254

::

world outside is changing rapidly.

255

::

That's what the data implies.

256

::

The core pillars of influence haven't really shifted.

257

::

So here is the provocative thought we want to leave you with.

258

::

If AI can now diagnose exactly which parts of our legal system are the most

259

::

complex, the most redundant or outdated.

260

::

Mm hmm.

261

::

If it can give us that perfect clarity, then the important question is no longer

262

::

can we achieve simplicity and fairness in our laws.

263

::

It's whether the political will exist to actually use this newfound evidence

264

::

to modernize the rules.

265

::

A profound question to watch unfold.

Video

More from YouTube