
The Neuron: AI Explained · 2026-07-01 · 53 min
Key moments - from our scoring
Substance score
55 / 100
Five dimensions, 20 points each
Tax AI, co-developed by OpenAI's forward-deployed engineers John Dewashage and Arthur Fernandez with Thrive Holdings, addresses a critical pain point in accounting: the eight-hour manual process of extracting and reconciling data from messy client documents (PDFs, spreadsheets, handwritten notes) before submission to tax software. The platform automates document classification, data extraction, and field population for complex 1040 and 1041 returns while keeping human experts in the review loop - practitioners upload source files, the system extracts and justifies findings with full audit trails, and accountants review before filing. The core innovation is a self-improvement mechanism where expert corrections become evaluation signals: when tax preparers fix misclassified documents or incorrect values, the system uses those interactions (not raw logs) to refine Codex's instruction set and edge-case handling. The architecture captures 'macro evals' - complete user steering interactions - rather than micro predictions, enabling the model to hill-climb on genuine errors. Codex has improved its ability to recognize knowledge gaps (especially since version 5.2), and the team observed the model eventually discarding hand-crafted skills it outgrew. Thrive's thesis is that AI transformation in professional services works best when driven from inside by domain experts, not imposed by engineers - practitioners shape product automation through daily use rather than engineering dictates.
Tax AI automates the extraction and classification of data from multiple document types (PDFs, spreadsheets, images, handwritten notes) and populates tax software fields for complex returns like Schedule C, E, and A, while providing full justification trails so practitioners can review and correct before filing - reducing what typically takes eight hours into a faster, more auditable process.
When tax preparers correct extracted values, misclassified documents, or file splits within the product interface, those interactions are captured as evaluation signals that guide Codex to improve its instruction set and edge-case handling for future similar documents - avoiding the need to manually log traces or engineer every fix.
The model began fetching and combining prior-year data on its own in ways the hand-crafted skill didn't allow, proving more effective, so the system proposed removing or updating the instruction - demonstrating that as Codex improved, it no longer needed human-designed workarounds.
Thrive Holdings acquired accounting firms specifically to drive long-term AI transformation from inside these verticals; the seasonal crunch (huge backlogs before tax deadlines) created an urgent need and a window to deploy and gather feedback quickly during peak tax season (January-March).
Simple ground-truth-versus-prediction metrics don't work because errors are hard to trace to their source - capturing full agent traces gives too much noise for the model to derive useful context - so the team built the product to capture 'macro evals' focused on the user's steering interactions that matter.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains some genuinely useful architectural ideas - capturing macro user-journey evals rather than raw traces or simple ground-truth/prediction pairs, and the narrow-coverage-high-accuracy deployment strategy - but the insight-per-minute rate is severely diluted by inarticulate delivery, constant filler language, and repetitive circling of the same points without adding depth.
instead of the micro but more of like the macro eval. So getting the whole user journey on what matters.
at some point the model wasn't even using the skill anymore. It was basically fetch that data by itself and combining it with other data in a way that the skill wasn't allowing it to do or wasn't designed to.
The empirical finding that a deployed model can autonomously propose removal of its own skills is a genuine and fresh observation, and the macro-eval framing is a non-obvious design choice; but the overarching advice (start with evals, work closely with domain experts, deploy narrow but accurate) is standard applied-AI product thinking that circulates widely.
at some point the model wasn't even using the skill anymore. It was basically fetch that data by itself and combining it with other data in a way that the skill wasn't allowing it to do or wasn't designed to.
you need to be able to measure exactly what you're working against. So basically building the evals and they are very specific to uh, the domain you're working on.
John and Arthur are legitimate practitioners - OpenAI forward-deployed engineers who co-built Tax AI and ran it through a live tax season, giving them real first-hand credibility; however, they are mid-level engineers rather than senior decision-makers, and their communication is frequently muddled, which limits how much value their genuine experience actually transfers.
I see you guys did like 7,000 returns, processed like a third of prep time, saved roughly 97% draft accuracy
John and I were not even based out of the US so we don't even file taxes in the US
The key metrics - 7,000 returns, 97% draft accuracy, one-third prep time reduction, eight hours to thirty minutes per complex return - are present but were largely surfaced by the host reading the blog post rather than volunteered by the guests themselves; the technical discussion relies heavily on vague language rather than named firms, concrete error rates by form type, or specific cost data.
I see you guys did like 7,000 returns, processed like a third of prep time, saved roughly 97% draft accuracy
sometimes you know, for example a tax expert that instead of, you know, spending eight hours by manually reviewing the files and you know, for example switching to like reducing it to 30 minutes on the platform
The hosts ask structurally sound questions about system-vs-practitioner task ownership, signal-vs-noise discrimination, and domain-applicability conditions for the agent loop, but they consistently accept vague or circular answers without follow-up probing, and never interrogate how the 97% accuracy figure is measured or where it breaks down.
what parts of that work is the system doing versus what the practitioner still owns?
How does it separate? And I guess this is where the human comes in as well. But, but I'm curious as to separating between, you know, a true system error of sorts and normal workflow noise
Computed from the transcript - who did the talking, and the words that came up most.
OpenAI and Thrive Holdings built Tax AI, a Codex-powered agent that helps prepare complex tax returns while preserving evidence for accountant review. In this episode, Corey and Grant talk with OpenAI’s John de Wasseige and Arthur Fernandes Araujo about how expert corrections become structured signals, how Codex turns repeated failures into evals and scoped engineering tasks, and why the best AI deployments still need humans close to the work. They also dig into what this pattern could mean for bookkeeping, audits, IT help desks, and other expert workflows where the system can measure what “right” looks like. Relevant links: OpenAI Tax AI case study: OpenAI Codex: Harness engineering: Thrive Holdings: Crete:
Transcribed and scored by The B2B Podcast Index.
Speaker A: So basically right now how it works is that the practitioners start with basically grouping all of the uh, data sources that they need to be analyzed. So as Artur mentioned this would be for example multiples and like huge quantities of like PDF documents. It could be Excels, it could be image and of climate nodes of clients. The model M definitely gets better at identifying when it doesn't know something but basically helping it going, you know knowing when it is good or not I think makes me think about you know, getting you know, good way of measuring what is true.
Speaker B: Basically I think what is interesting about this is that the uh, that part of the product that matters a lot to the practitioners is basically being driven by their use and not by engineer dictating how that works. The, the structure of the text engine helps us a little bit get provide initial structure and comparison compared to if you didn't have this, you just had like the end documents.
Speaker C: Welcome humans to the Neuron AI explained. I'm Corey Knowles and I'm here as always joined today by the one and only Grant Harvey. How are you Grant?
Speaker D: I'm good. I'm surprised you didn't throw me a curveball today.
Speaker C: Today the curveball is that I didn't throw you a curveball.
Speaker D: You've done that curveball before so it's a double curveball.
Speaker C: Oh, oh, I've got you twice.
Speaker D: Well I'm good. Corey, how are you?
Speaker C: I'm good, I'm good. I understand we're going to talk about taxes today.
Speaker D: So today we're joined by OpenAI's John Dewashage and Arthur Fernandez, uh forward deployed engineers to Thrive Capital who helped build the joint Venture Tax AI with Codex. So people are familiar with OpenAI? No, Codex Tax AI was apparently built for real accounting workflows. So we're talking about messy client documents, we're talking about source evidence talk tax software mappings and human review. And the big idea is using a self improvement loop where expert corrections become evals. Um, Codex investigates the relevant traces and code paths and engineers review bounded fixes before anything ships. It's very, very exciting.
Speaker C: John Arthur, welcome to the Neuron. We're so excited to have you guys.
Speaker B: We're excited to be here.
Speaker A: Yeah, thanks for having us.
Speaker C: Well I guess let's start here for listeners who have never seen Tax AI, can you give us a what it is kind of walkthrough?
Speaker A: Yeah, sure. So basically Tax AI to give a bit of context is a platform that we uh, that Thrive holdings uh, co developed with uh, us. And so basically uh, the context is that you have Thrive holdings which is a subsidiary of uh, Thrive Capital. And what they are doing is that they are uh, through an intermediary company they are buying like roll ups basically. And so they are buying companies in different, like in different verticals. And so one of them is taxes, the other one is uh, IT services. And basically we are working on the taxi side. And so the idea is that through those companies they are, they are helping them by creating and developing a ah, platform. And so what tax AI does it basically it's a platform that allows you, that helps the taxpayers to accelerate the tax uh parsing from the data to basically submitting into the tax engine but also adding a lot of other features uh on the side. And it's basically bringing AI helping the taxpayers do their job. And we co developed it with them and helped them on the AI part of it.
Speaker D: Very cool, very cool. So it's a, I understand that it's a product more for um, like let's say like accounting firms and less for individuals. Right? Um, what, what was the impetus or why, why did you want to create a product for, for them specifically? And, and how did you approach it?
Speaker B: So holdings specifically, uh, like the way that they set up the business as John explained is that they are working with these verticals and they have a thesis that they are places uh so they, they don't hold a short term holdings like with these uh, with these assets. So they think those are industries that benefit from long term AI transformation. So that's why they started with like accounting and IT services. And a strong belief from holdings is that in general like AI, not AI uh transformation like technological transformations, they usually hit the industries from the outside and they think is uh. These specific verticals are ones where this transformation can be driven from the inside out. So they had this strong belief of essentially driving this transformation from the experts. So basically co building a product with these experts. And OpenAI has, has a partnership with holdings where we also hold some stake on uh, on, on, on these assets. So essentially we are FDs helping them uh build solutions around this space. And this problem in particular was the first one we started working on with holdings. So as, as kind of like any problem like you come in there like what, what should we build? And the reason why we picked preparation is we identified this as one of the biggest pain points that uh, uh the, the experts uh surface to us. So if you think about accounting uh with the perspective of like uh, the accounting firms uh, it's very seasonal and they get a Huge backlog in the weeks just before the deadline. Yeah, and a huge, A huge problem for them is basically my fault by the way. All of our thoughts and uh, and the problem is a lot of the work is done, for example, uh, on the more complicated returns, on basically going through RAW files with all sorts of different formats, from more structured PDFs to like spreadsheets, handwritten notes. Sometimes you have to go and pull formal information from the client and you have to get through this with a deadline. Sometimes it takes like eight hours to go through these documents and just input this data into the tax engine. It requires reconciling information from different documents and things like this. So the problem we defined that, uh, had a lot of value and around the time we started working with them, which was December, we also had a huge opportunity to capture a lot of the value if we deployed it system in production super quickly that we could leverage, like we, we could get a feedback like January, February, March is something that's super useful to the firm. So, and we talked about this a bit in the blog post around the overall impact. Uh, but in general, like, we decided to start with preparation because it was a huge pain point for the experts. So they were definitely on our side to like start working with us around like, whatever can help us like reduce the gap here and the time. Um, yeah, makes sense.
Speaker D: And it's cool that you tried to get it in right before they're taxed. Yeah, yeah, that's great. Got uh, it.
Speaker C: So when you say, uh, tax AI prepares complex 1040 and 1041 returns, what parts of that work is the system doing versus what the practitioner still owns?
Speaker A: Yeah, so basically right now how it works is that the uh, practitioners start with basically grouping all of the uh, data sources that they need to be analyzed. So as Artur mentioned, this would be for example multiples and like huge quantities of like PDF document. It could be Excels, it could be image and of climate notes of clients
Speaker C: or all of them, probably.
Speaker A: Yeah, all of them. Exactly. And for, you know, depending on who is sending it, you know, for some it might be some very more clean PDFs. For some it might be way more, you know, like just screenshots and client notes. And so you get, you know, uh, like you go across the spectrum of all of them and this makes it, you know, like you can have one process that basically automatically analyzes everything. Okay. And so it starts there. They upload all of this data in the platform. Uh, there is some also work that, you know, in some cases you can directly connect them from the email and upload them on the, on the platform. But basically the key part is once it's there, what it does is so there is some. The software is running and then it extracts. It's going to first split the files, uh, in order to identify and to classify them. Because you might have everything that it's in one big 200 page document, for example. So you need to basically classify it and know, okay, this corresponds to W2, this corresponds to K1. So this part is also very important. Then they see the split and then they are able to see the data that has been extracted immediately. And so you would see here the split of a W2 of a K1 of a schedule E. And you're also able to see a justification of why this data, uh, is extracted basically. So you're able to see, okay, uh, this value comes from this Excel on this page, from this cell, but also is reconciled with that image, that PDF and that data coming from there. Uh, and so this is also very important for the taxpayers because it's definitely one point in which you cannot just blindly, and you don't want to blindly trust the system. And also you want to make sure that you don't end up spending more time on actually needing to find why the AI got to that solution. So what you need to do is make sure that the justification is very correct and very accurate, uh, and that you're able to basically see it in details and go back to the files to understand what was there. So what the taxpayers do there is that they're able to review what has been extracted. And so a lot of their time is actually spent on focusing more on the review process and making sure that the more complex fields are correct compared to basically the more simple ones which they would still usually need to spend that time. This would be, is correct all of the time by the AI, basically. And so they are able to focus on those harder ones and then they review this and so they spend a lot more, much more time on the review process and then they are able to make the submission on the tax engine.
Speaker D: How does Codex keep it straight? Uh, like all of the facts. Like, is it using some sort of like OCR model? Is it, um, like doing like loops on the back end, like to be like, okay, I've already. Does it have like a running list of everything it's already checked against? Like, I'm just so fascinated how that, how that works.
Speaker B: Uh, yeah. So basically what John described, and I think Codex is very good actually at following Instructions in general. So what John described is sort of a workflow right in terms of like how you process a file. And so in general uh, I think to summarize like you, you essentially have like a set of steps and you have a set of uh, durable artifacts that you're following through these steps. Because also some of these steps might also like fail and you might need like retries and things like this. So you need the durable execution. But Codex is good at following instructions and it can basically drive like, like new facts like new files coming in and getting like stuff into, into that like end step around basically the, the reviewer ready. So it's basically a system that works through like, like uh, churning through these durable artifacts. And the harness around it is essentially the, the set of instructions that were, that were catered and when we're talking about self improving in the article is essentially that harness the set of instructions skills that like everybody's building their own skills uh as well like these days. So it's, it's the set of things that you're basically giving the model that helps it for example deal with certain edge cases. So let's say uh, in the example that John uh gave that uh the like a few tax preparers are having to override a specific field. So basically the system is starting to identify basically a pattern here that like we're always extracting information incorrectly for a specific field in the tax engine or maybe it's related to a specific type of document that we classify that we perform more uh, like more poorly. So those are kind of like the things that we try to uh, get the self improvement loop to contribute changes back to is essentially like the mechanism around handling these edge cases and also making sure that like uh, because we can measure things well like making sure that like we don't regress like around what we supported before.
Speaker C: What are the. What's kind of the hardest real world messiness you found to handle? And is there anything that was easier than you expected?
Speaker A: So one thing that's definitely one of the hardest I uh, would say so I would say on the personal side you definitely have for example things like Schedule E, Schedule C and schedule A's. Those are I would say are usually very complex because you need to reconcile many different files together. But then also you need to sometimes in those files you don't have the full context itself. And so usually you would need to look at prior year context. And for example if you start looking at the preassociation or amortization like this is also you Know adds a layer of like understanding the tax workflow. And so for example in those cases it's also very useful to give context, uh, on taxes. So for example, you know, if you give IRS documentation to the loop that extracting it then it's also very useful. And so this is on the personal tax side and then you have also the entity tax side which is you know, kind of separate and works differently. And there, there is also like a huge amount of messiness because you need to start looking into trial balance. And this means that you know, usually you don't have two same companies which are doing things in the same way. And it can be also even more messy because the quantity of data or the number of rows can be in the even larger uh.
Speaker C: Oh wow.
Speaker B: Yeah.
Speaker D: Oh, did you find, were you surprised with how Codex was able to handle a lot of this? Like, were you like oh, it actually can do this better than we even thought it could or what did it take a lot of that tooling and self improvement loop to finally get it to a place where you're like yes, it can do this reliably.
Speaker B: I think what's interesting here as well is the pace of some of the models that we are deploying and also um, Codex app itself and a lot of the stuff, some of the concepts if you think about the app was released in January if I'm not mistaken. And then we have multi agent and we have more things that enable you to codecs now can manage like different threads. So you have all these different capabilities that enable you to handle a lot of the orchestration, have more like independent reviews and also like other processes that you can incorporate which in the past you would need to like build by hand. And also because in January I think we had the 5.2 and now we're like, now we're like 5.5 and like the pace of changes and intelligence that is added to the models is frankly like quite impressive. So I think we could only get to um, the point we got on when we talk about in the blog post is the models is at a state that given a uh, well bound task that is measurable it can do that reasonably well. So I think it's around. The main problem is around now I guess defining the right objective and there will still be instances where the model potentially is not going to be able to automatically uh, climb that loop because it's something that might be potentially ambiguous still and requires like intervention. Ah, and going through the original document and seeing like okay, like this is a very different Format that we haven't imagined like, or this is like completely misclassified. The documents lack some capability around the visual language model side of things. So there are still like some capabilities around that space where it was still going to be a source of errors or uh, a source of problems. The, the model is not going to be able to automatically heal climb. But with every generation we are seeing like massive gains in intelligence. And with every new generation as well, like some of the things where you had to provide a lot of instructions on, it reduces your need to do some of these things. So we also think a lot about basically like what is like some of the things that you can shed away like, like uh, that uh, you don't need anymore because the model is going to have that in distribution or is going to be smart enough about not making like not requiring that specific uh, prompt. It might even have like a more creative way around resolving that task than what you're right. Oh wow.
Speaker A: I think a story on what Arthur just described that one thing that we noticed, for example around I think it was February, is that at some point we had like basically in the extraction loop we were giving a skill to the model to basically look for uh, you know, previous like frequency data of older text forms. And what we observed is that at some point the model wasn't even using the skill anymore. It was basically fetch that data by itself and combining it with other data in a way that the skill wasn't allowing it to do or wasn't designed to. And so, you know, propose it basically. Okay, I'm, um, not, you know, you can. It ends up basically creating a pull request to remove it or to you know, like, in this case it like added some things into it, which is interesting as well. And so you want your system to be able to, you know, propose changes that, you know, keeps its interest, like, keeps its focus and allows you to do more of it.
Speaker D: Yeah, that's a good point. Because even my own skill creation that I do sometimes I'll notice like, you know, maybe there's a model change or something or I start to shift what I'm looking for from the skill and the skill can sometimes be holding me back. Right. So it's cool that the model knows like, hey, I need to move past this and I can do this a better way that's building on that.
Speaker C: Something that I think is a good call out here. And I think this is not possible. If the models haven't improved to a point they can do this is that they do. And I would Say this is probably sense 5, 2 forward even is the ability to recognize, oh, I can't do this and come tell me I can't do that instead of just winging it and doing it anyway. And uh, there's a lot less of that than there used to be. Did that play a big role in making this work? Because I mean I feel like with the loops we're talking about, with the practitioner in place, its ability to recognize, hey, this is a problem, you need to check this out. Has to have played a big role.
Speaker A: I think that's uh, totally an important point. And one thing that from the experience and from what we have seen in the past months, one thing that helps uh, on this and the model definitely gets better at identifying when it doesn't know something, but basically helping it knowing when it is good or not I think makes me think about getting a good way of measuring what is true. Basically giving the model M that ability is super important. And so one thing that we spend quite a lot of time making sure that was in the product and this I think is more on the architectural side of how you define your product, uh, which the AI ah, doesn't always do great, uh, this is. Yet it's basically you want to make sure that when the user is going to spend time on the product and fix something. So the thing we described was for example correcting a value or, or saying it was written in the bad field or you know, like splitting it, saying that the file had been incorrectly split at a point and should be split another place. Um, capturing those things and letting the user do it in a way that the output of what they are doing can be used afterwards to guide the model and say, okay, hey look, if you thought that you know, the split was there or the correct value was this, this is untrue. And use all that data from the user. And so, you know, it's not exactly about recording everything that the user does. It's more about knowing exactly where you want the user, human, human input to, to be and the expert input to be such that afterwards you can help the model know itself, you know, find out by itself that it was wrong on those places.
Speaker B: So I, I think a lot of people are also thinking about uh, what's the information you track around, like agent traces for example. Ah, so I think in one end of the spectrum you have a building in EVA with just the ground truth and the prediction. But what we realized in our space, related to what John said is that this is super messy and it's hard for the model to reconcile. Like you, you can, you can measure where the errors appeared, but what exactly is the, is the source of this thing. And then on the other end of the spectrum you can capture let's say the agent traces like the raw application traces. But this is a lot of information for the model to derive context on. So I think one of the things uh, when we designed the product that John mentioned is trying to capture more of the, instead of the micro but more of like the macro eval. So getting the whole user journey on what matters. Uh, so if you build your product around capturing that so like the interactions where the user has to steer the same way that you steer your codex session sometimes to get it on the right track. So getting the right like steering uh, is basically how you can basically build the user journey for the relevant signals. So this is way more useful data for the model to do hill climb on the errors essentially in triage.
Speaker C: And I guess ideally you need this to be a tool that uh, any accountant is capable of using. Not just a computer engineer is part of a tricky element I assume.
Speaker B: Uh, no. Yeah, for sure. So I think one thing they're very proud about this project in particular is that John and I, Even though we're FDES here, OpenAI, we come from a uh, product, uh engineering background. So we, we built products before like the, the folks that Thrive holdings they are building. We did that in the past and in the traditional software engineering lifecycle it's usually the engineer's job or the product manager to structure things in a way that you collect feedback, you triage and then the engineers are going through. Maybe if it's like more driven by logs and things like this that basically fixing issues. But I think what is interesting about this is that uh, that part of the product that matters a lot to the practitioners is basically being driven by their use and not by engineer dictating how that works. Of course the engineer gonna set up like the architecture like how maybe some of the UX thing things are gonna work in the front end. But basically the automation that immediately impacts the time they're gonna be spending on the review is being driven by the more they use the product essentially. Which is. Which I think is like an interesting anecdote to on. On like on like deeply integrating with, with the experts. Like I think as fd we have a very humbling experience in general because we come to these deployments like John and I were not even based out of the US so we don't even file taxes in the US So instead of become uh, like tax experts, like you need to basically give uh, the people that know what they're doing, like uh, you should not trust us to do your taxes, please don't do that. But actually give these people the power to shape the product uh, in the way that fits their use case.
Speaker D: It's sort of like the next generation or next evolution of that feedback loop with the software development cycle where you're almost letting the product itself get the feedback from the user. Because if I'm understanding this correctly, at the point where it's at now, the people who are using it, the accountants at these firms, they are uh, giving it feedback and it's improving kind of autonomously. Right? Is that more or less what's happening at this point?
Speaker A: So uh, I would say like, so they are giving feedback and the feedback is triage. So like when there is a new feature request, it's definitely triage across the others because on the other spectrum you don't want to directly authorize that. You know, when someone asks for a request it automatically becomes, becomes a new feature. You know, that could be interesting for that very specific person, but that actually could be a pain for the rest of the users using it. So this is also something else. So there is kind of a human review at some point in that early feedback phase afterwards in that iteration loop. This one, you know, can be self improving, but on the feedback side from the user this is definitely important because one thing that's really interesting here is that in one way you want to give, you know, to make the product as smooth and as easy to use as, you know, Artur was describing before. But on the other end you don't want to allow any kind of feature to be added because you know, in some way if you create this, you know, you have a, uh, tax expert that instead of, you know, spending eight hours by manually reviewing the files and you know, for example switching to like reducing it to 30 minutes on the platform, if you start by adding some tools that, you know, for example a game to play while the data is being extracted, or a chatbot to click without any intention, for example, it might mean that the person, instead of doing 30 minutes a day, will just end up spending two hours for something that helpful as well. So I think this idea of making sure that you stay focused on what is actually truly useful for the user and for the end task, uh, is quite important. And yeah, you basically don't want them
Speaker D: to spend eight hours just messing with the tool instead of actually like being able to Serve more clients.
Speaker C: That makes sense.
Speaker B: Yeah.
Speaker C: How does it separate? And I guess this is where the human comes in as well. But, but I'm curious as to separating between, you know, a true system error of sorts and normal workflow noise, like just the difference from tax return. I say return from tax filing to tax filing. I assume big ones aren't getting returns easily. But uh, I'm curious because there's a certain amount of uniqueness to every filing. Even though the forms are the same, how does it recognize this return is different from this return, but not in a way that means something is broken, like the uniqueness. Yeah, I hope that makes sense.
Speaker B: No, yeah, it definitely does. So I think in our problem space in particular, there are a few things that give a bit of structure that helps sift through these different things. Uh, I mean there's still a, uh, certain amount of ambiguity which might lead to not the right. As I mentioned before around like maybe it's a completely new type of document that was misclassified and things like this. But there are a few things that basically make some of the focus for the test. So classification is one of them. Also the mechanism at which we deploy this product, we don't support the entire surface area of like filing on all of the states and uh, all of the aspects of federal tax and entity tax. So the way we deployed the product was very targeted. So we started with even with the simplest of forms. And of course when you do this initially the probability of this being more correct is higher. So that kind of gets uh, us the right shape for understanding exactly how to capture things. And we don't have the loop at that point. But also the tax engine gives us some guardrails around the fields that are filled in there. So the tax engine that the firms use, it kind of already encodes a lot of opinions on how you put in a value. So it's different. For example, if you ask Codex around your return in general, Codex is going to be taking care of parsing the file and also doing the calculations and things like this. But the text engine, for good and for bad, it provides a uh, lot of guardrails around exactly what to expect. Why is that? Field. Field in a specific way. And that gives us basically a way to more strongly measure errors. So for example, the error is going to be anchored on a few fields, for example. And then we start seeing the repetition of like a few tax returns like are making mistakes on like these specific fields. It does create some problems around grading the task correctly because some of the fields Might have more of a human preference and like uh, like some of the fields, like an accounting firm might do that slightly different than the other. So the way that we encode some of these things is by going through these initial interactions with the uh, with like the practitioners. So we get some of these things encoded in our grader. So it's, we only get to the more like self improving loop once we get some of these things right.
Speaker D: Uh, I guess if I understand this correctly it would be like the tool that the companies use. It already kind of has details of like this is how we report this type of revenue or this type of uh, situation here. And it's able to like parse that to some degree. So it's like already there's like this is how our company does it. Tell us a little bit about what that is like and if this process of what you're doing now would be possible without forward deployed engineers or how integral that is to being able to work this closely with the clients.
Speaker A: I think it's very interesting because it's uh, a combination of someone who uh, is a software engineer but also is potentially a product manager, uh, has concept of architecture and is able to deeply ingrained with ah, the, with the customer. And so instead of you know, having the people who actually go with the customer separated, fully separated from the people who are designing the software, the idea is that by being very close to the customer you're able to understand the intricacies of, you know, what you need to crack to basically allow a new workflow to be, to be solvable. And so in the case of OpenAI, what is really interesting is that you are very close to research or very close to engineering in order to shape how the models are going to improve, know where they are currently lacking in terms of results and where they are not able to do some tasks. You're also close to engineering because you work directly on the product that engineering is working on. You help them go in certain direction, you add features when there is a need to, but then also you know, you're making sure that what the customer is trying to achieve, you're able to get them through. And so you work on very specific problems which are unsolvable and which basically you try and go to see from 0 to 1 if it's actually solvable or not.
Speaker C: I'd say that shows up too from just the proof points in this project itself. I see you guys did like 7,000 returns, processed like a third of prep time, saved roughly 97% draft accuracy. That's a huge number. And I wonder where is that the most meaningful to a practitioner and where is it most to an engineer?
Speaker B: Do you mean like uh, in terms of the value we give the engineers that we work with and the practitioners?
Speaker C: Yeah, I think the way I meant that and I realized it was worded really poorly is like, is that high accuracy the most important point to the end user? I assume when it comes down to it, 97% strong. But I assume there's always a desire to go farther.
Speaker B: Uh, yeah. So I think that is an interesting number to anchor on because if you think about the, from the perspective of a practitioner and also we talked about the iterative uh, like basically expanding the perimeter of the product more strategically so starting very small. And I think this is very important because if we support a lot of coverage, so we covered a hundred percent, but our accuracy was very poor, this is a complete failure, right, because they will be spending a lot of their time, they will have no trust on the product. But if we deploy something where it will cover some of their work like that they will have to do like things that we classify that we can support. Like you still have to go through this manually and that's completely fine. But if you support those with uh, with strong accuracy, I think there's a huge win. So we really focus on, on trying to expand the parameters strategically and not really surfacing a lot of errors for people to fix because we imagine this would just generate a lot of friction to, to, to, to practitioners. And we think that in problems that you, you can really measure the outcome. This is a good way to deploy uh, like in, in a lot of fields, sometimes like you build for example like a go data set. But when you're, when you're working through this with, with an expert, you're basically making them commit extra time to basically build this with you. But if you can deploy something that is immediately valuable and can basically give them like a uh, like a, basically it helps them with something that is like lower complexity or mid complexity, there's already a huge win. So I don't think we should ever track like getting to 100% because this is a world that's always going to have a lot of nuance and some ambiguity. So the point that John made uh, like some time ago on surfacing the, the citations back to like what the numbers mean, we think that's always going to be an important capability of the product because in any, to understand precisely, I would say it's kind of a model receipt. It gives you a Receipt on exactly like here is your bill, it costs you a hundred dollars but here is the breakdown. So it's, I think it works in a similar way. So you can understand exactly where, as
Speaker D: you approach this process, um, you know, from an engineering standpoint, from a product standpoint, how did you think about the cognitive load of the expert? Like where, how did you balance like how much they could actually process ah at a given time? Because I imagine with some of these accounts that are more complex and you're offloading a good amount of the thinking to Codex in this case or Tax AI like how do you still keep them in the loop so that they can still keep track of all of it in their head?
Speaker A: Yeah, so basically what we, in order to kind of you know, make sure that you, yeah, you need to calibrate the product to make such that what you know, the preparer does and see, you know, it doesn't overload them. But then like my feeling is that this is a lot about also calibration and making sure that when you iterate and when you develop the tool you stay very close to them. So um, here uh, the Thrive team and us, we also spend quite some time with the taxpayer and close to them. So I think there is a first point of making sure that you understand very well how they are doing things currently what they are to process on a day to day basis. And there is this thing where you don't want to disrupt totally the new flow by saying hey look, just drop everything in there and then you know, you trust it and then you submit it. You know like could be way higher, they could process way many packages but it doesn't work right. And you want, you know, you want to make sure that you keep um, the expertise, uh, where it is in you know like account on that uh, old expertise and so staying close to them by you know, making sure that when the, you know like having this closed loop iteration and you know delivering and shipping quite quickly to make sure that every time you do something you're not disrupting their flow and it's still in that capacity and it makes their expertise and their knowledge to go into the most interesting tasks basically is what is definitely important. Um, also maybe taking a step back on the kind of achieving. So basically when uh, we started this workflow there was not a goal of trying to reach 98% of accuracy for those packages. I think what's is it's like how far can we push the taxes workflow and how best can we make it across individual taxes entity Taxes and there is no real limits. The same way you would say it's done for models in evals where basically you might say that two years ago you had some way more simplistic eval compared to now where it was okay, you need to solve all those math problems, you need to solve those computer science problem. Now the evals have been saturated and so you focus on much more long running tasks. I would say there is you know, some meta point that's a bit similar here where it's you know, you have at some point solved or almost solved the uh, W2 form. Then you start focusing on the more complex one. But then you know, once you have solved most of those then you know, how can you help even out of the world, you know, and how can you take a step back and you know, make sure that you're solving long running tasks even more?
Speaker D: It's all about changing the frame. Like uh, there was a great uh, article that Dan Shipper released that was about like uh, you know evals are all about changing the frame. Once you change the frame you like start from zero again and it's like now you have a whole new thing you can work on. I have a question for both of you. So after working on this project, in your personal opinion, do you think taxes in general are a verifiable uh, solvable like quantifiable task or is it more of an art? I guess I'm asking is there more art to taxes now that you've been working on this? Or, or is it more like something that like at some point we'll be able to fully solve?
Speaker B: I think for our problem in particular we were solving uh a bounded part of it which is essentially just the data entry part of things. Uh there is something uh, that I imagine there is more nuance around ambiguity around tax rules uh, which uh, which I think like uh. That's probably a space which has, which requires more of the expertise and also some of the expertise in our case is already encoded in the, in the tax engine in a way which provides some of the structure and some of uh, some of the things related to this. Uh, but I think the, I think the art part which is also uh, what, what uh holdings is trying to give the companies by giving them more time uh to, to spend on strategic work is how can you provide more tailored feedback throughout the year. For example so, so the clients uh, uh, can basically better uh, like based on prior information how the next year is going to go. And I think that's an area that Has a lot of ambiguity, a lot of context dependent things where uh, that's where I think the art is. And that's probably the more strategic work that the moment you unlock experts from doing some of the things that they definitely don't want to do around just churning through documents, they can focus more on that of their time around that strategic work. And that's where the expertise shows the client management, the previous client relationship. I think that's where it shows and that's, that's the human part, that's the art part I would say, you know,
Speaker C: in this uh, I remember OpenAI and thrive in the blog post I believe it was mentioning that this is a pattern that could apply to bookkeeping or audits or IT help desks or even other operational work throws. What workflows? What has to be workflows. Sorry. Or other operational workflows. What has to be true about a domain for this kind of agent loop to work.
Speaker A: Yeah, so the interesting thing is that, so right now what we did with Thrive currently is uh, focused on taxes. But so Thrive holdings has other uh, verticals on which you're focusing on and another one is IT services. So basically where you want to solve IT ticket or help engineers uh, solve the IT ticketing that they have. More generally the loop that we have is usable on you know, I would say many kinds of problems and you know, currently I wouldn't see a limit. I would say whenever you have some experts, uh, you know that know the field or that are doing something when you have a product that is being developed where you are actually able to track, basically track what the true value is or what the correct value is, is a perfect setup for this loop and it's great. So in our setup we made it specific to the taxes overflow but it's totally expendable because like things are centered around codecs, uh, the way you structure basically your software and all the uh, components and things that are saved throughout time, this can be kind of replaced with the ground truth, the evals. The evals is something which can also be generic. And so what you end up seeing like what you want at the end is to, you want to make sure that there is some users using a product where the ground tools that they have can be used beneficially in the long um, term. And this kind of setup I would say applies to an infinite possibilities of sectors because you know, anyone who is basically working part time from their computer doing something can have, you know, on a certain product can have the skill base there.
Speaker D: So how Would you, if you were advising a company to try to build its first like domain agent, kind of like following the same roadmap, what would be the first thing that you would tell them to do or set them set up so that they can get that like kind of same process going?
Speaker B: I think for us and OpenAI and maybe for other people doing these applied applications it might be uh, it might be pretty straightforward what I'm going to say. But um, you need to be able to measure exactly what you're working against. So basically building the evals and they are very specific to uh, the domain you're working on. So I think that's the first uh, place to start. So other is when we go to a new project, we have the scoping and all of this, but we always start from evaluating first uh because it's basically the, the first building block that you need to leverage. It also helps you think about exactly like what is the right hill to climb. Uh, so I think, and that's very domain specific, so that's the first thing you have to answer, I think.
Speaker D: How do you approach the evals? Is it like a goal? Do you have like a golden set where you're like this is what a perfect version of this task looks like or is it a little bit more broad? Like what's your advice on just like how to set up a good email?
Speaker B: So I think uh, a gold set is definitely very helpful especially in the initial discovery. But in our domain we have the finalized returns done by the experts. So in a way we already have the gold set. There is some complexity as I mentioned, around making sure that we remove some of the things that might be file and different places and things like this. But the structure of the tax engine helps us a little bit, uh, get, provide initial structure and comparison compared to if you didn't have this, you just had like the end documents. Uh, but I would say like depending on the domain you might need to go the direction of doing more of the manual curation with an expert. You might have no other way to get through it. Or in our cases like basically you have a lot of data being generated on the field already that you can leverage.
Speaker C: Much like in mathematics recently where you all discovered, I say you all not you specifically discovered an error in a known problem. Uh, I suspect we'll see that approach at some point too where it's finding problems in human produced returns M and we'll probably see that in every field I would assume at some point. Do you have any thoughts on AI and this is maybe a little off topic, but do you have any thoughts on AI in that role in catching mistakes we've made versus just us trying to catch mistakes from it, which has been the approach so far?
Speaker A: Yeah, I think it's a, uh, very like the mass problem you mentioned is a very interesting example because it relates to something which we, in our case, for the tax case, we have also observed is that for some of the cases, as Artur mentioned, when you're building your golden set, you basically are using the. All the data that has been submitted from last year to the irs. And so the cool part is, you know, this is true and. Or this is. No, it should be true. It's supposed to be true. Right?
Speaker D: Yeah, you might, you might be able to find some flaws.
Speaker B: Yeah, accepted this. Yeah.
Speaker D: Does IRS have a whistleblower fee? You might want to hit them.
Speaker A: And so, yeah, this is definitely, uh, a great luxury because, you know, like, not, as Arthur mentioned, like, not every use case does have this. And so in our cases, what we sometimes saw is that the model, like, uh, this loop, basically the software was extracting some fields and so we would see some prediction when we were rerunning the device where, when you like, what you often need to do is like deep dive into a specific thing that has been extracted to understand and try to see, okay, is it good or not to understand a bit the errors. And so we noticed that it had extracted some data and that it was actually true, but it was marking it as false because in the ground truth from the irs, it was that this value. And so you're diving into this. We realized that the actual submission that had been done was off, you know, by almost, you know, not a lot. So it was not very important in the end. But it was a very interesting case of, hey, look, you know, the model found something that, you know, it's actually the real ground truth. And the prediction, like the thing that we thought was a grand tool was not. One thing that's interesting there is that our take on this is basically you have, you know, when you have an agentic loop that's running, it's going to be up and you know, in a software, it's going to be up with the same energy all the time. And so there is with humans, you can have, you know, energy and concentration and ability to do things that, you know, is, you know, can be more modular or, you know, you can have some lows. And so in the end there is this idea of you want to make sure that you concentrate what people are Best at in some specific case. So in the case of like a huge tax return, you want to, you know, if they are spending most of their time on the more complex stuff that's more interesting and maybe even more funny to do in some way. And you know, removing that part of, you know, just copy pasting a uh, W2 from a box to another one, then you, you know, you potentially can decrease those errors. And that's what we hope as well to see.
Speaker C: In this case it makes sense because in the AI discussion a thing that I know, especially people who aren't as micro analyzing it as we are, there's this temptation to believe that if a human's done it, that means it's perfect. And the truth is we are not perfect. Uh, we are very fallible. And I love the idea of a backstop or a safety net behind me to uh, to catch my own errors. I think that's, that's a uh, huge power that, that we overlook. And I love how this works. I think this is a really, really cool approach to a problem that everyone here dislikes. Like even, even for accountants, this is a tedious thing. Like it's, it's. And it makes a lot of sense because with such a complex tax code with so many variables that come into play in taxes themselves, that this feels like a thing that if it can knock this down flawlessly, there are a lot of other things it's going to do really well with just because also
Speaker B: I think this is one of the type of tasks and I think there is a bunch of tasks that are similar to this, which is like producing the extraction into the text engine requires a certain amount of effort. Reviewing requires a different amount of effort. So if you get sufficient accuracy on, on the things, one, one shotting like uh, and you have to fix a few things, reviewing is, is way less effort than you would. You would do that job yourself. So the overall accuracy of the system increases overall because a lot of the work, as John mentioned, they'll be consuming a lot of your attention, which is maybe the causes of some of the errors and also like fat fingering or doing like some of some of these mistakes.
Speaker D: Yeah, just the amount of stuff you have to review and then you get lost in the details and you miss like the most important thing because you were too busy. Like I had to double check all of the, you know, my new.
Speaker C: We're so distractible.
Speaker D: Yeah, well, I had two, two more things before we wrap up here. I'll make them quick. The first one is this almost makes Me think, like, I wonder if this will lead to sort of like a adversarial, uh, game, uh, where the IRS will have their own AI tax AI and then the firms will have their tax a. And then they'll be trying to one up each other. Kind of like how today's accountants sort of are trying to do that with the government. I wonder if you have any thoughts about how that might play out, which is interesting.
Speaker A: Yeah, that's very interesting. I actually never thought about that case. But, uh, yeah, what you could have with this is maybe you would end up having a conversation between someone's tax software and the irs, and each of the, uh, software AI could post a challenge, like a message on the channel saying, okay, this is true. No, this is true because of this. And then the other one saying, I agree what I thought was wrong. And then the other one, I think you need to have some, uh, kind of framework if you want this to be working. But in the future, you could have something like this. And I think it would definitely help solve some use cases where sometimes you might notice a mistake that was easy to fix. And if you would have such a direct interaction, you might fix it way earlier than having to spend way, much more time, months afterwards. Uh, I would think so. Yeah. This would be.
Speaker D: Yeah. Because a lot of people in the US Often say, well, why do we have to figure this out on our own right at the end? Why can't they just tell us how much we owe? They know the taxes code, you know, why are we trying to.
Speaker C: Because a lot of people do their own returns in the US Is a thing, too.
Speaker D: Yeah, exactly.
Speaker C: We're not. We're not tax law experts.
Speaker D: Yeah, exactly. So it would be kind of nice if you could just like, call them up and be like, hey, uh, I think this is correct. But is this correct? And your agent could just kind of chat it through.
Speaker C: Codex talks to their agent, and, uh. And I just find out when it's finished.
Speaker D: Yeah, yeah, exactly.
Speaker B: Uh, some of the problems as well are the different reporting mechanism. So your receipt, you'll be leveraging, like your accounting firm, uh, to do that. Like, I think for some, basically, maybe you have the same fact being reported, like, you're guaranteed to have the same fact, like your employer. But I think for some facts, it's slightly harder, like, uh, like charitable contributions and. And other things like this. So I think that's where, like, I think things get, like, more complicated around, like, uh, what are the facts the IRS holds and what are the facts that you hold and uh, making sure that those reconcile correctly. Yeah, yeah.
Speaker D: And then my last question to both of you is, you know, what do you Want to see OpenAI? How do you want to see OpenAI improve its models from here? What do you really want to see from the next generation of models? What could really improve this use case and really help you with what you're trying to do?
Speaker A: I mean, uh, I think for on this specific use case I think what we see generally is that the ability to do long running tasks definitely helps on a lot of other things because usually if you're able to do more long running tasks it means that the model is able to look up in some of the data. In our case it would be the evaluates the traces, uh, also being able to do some tool coding, uh, imagine at some point being able to do some part of the tax calculations, uh, itself in a, in a way better way. This would be interesting as well. I mean I think generally speaking like the more it becomes able to do things that are related to, you know, the economy in general is better because I think in taxes you have a lot of concepts which are human invented and you know, which are very potentially hard to grasp for uh, a model and so, you know, the more it understands the economy. And so I think we have for example evals that are well known on this such as GDP eval, uh, which are more broad. I think expanding into those domains means that you have taxes that improve directly but then also has consequences of other things improving. And so yeah, this is quite exciting and will be useful for the next model for it.
Speaker C: Absolutely will. John, Arthur, thank you so much for taking the time to join us today. This has been really interesting.
Speaker B: Thanks a lot. Uh, we're really excited to be here.
Speaker A: Thanks a lot, that's awesome.
Speaker D: Yeah, great to hear.
Speaker C: We'll share the link to uh, the blog post about Tax AI as well and make sure that it's there for anyone who wants to check that out and learn more and see what's going on. Please take a moment if you haven't yet to subscribe to the channel so you can check out all of our interviews with the people building and impacting AI every day. Also, please pop by the Neuron AI to subscribe to our daily newsletter. And that's all we have for today ladies and gentlemen, but we will be back soon with more farewell for now humans.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.