Intellectually Curious is a podcast by Mike Breault featuring AI-powered explorations across science, mathematics, philosophy, and personal growth. Each short-form episode is generated, refined, and published with the help of large language models—turning curiosity into an ongoing audio encyclopedia. Designed for anyone who loves learning, it offers quick dives into everything from combinatorics and cryptography to systems thinking and psychology.
Inspiration for this podcast:
"Muad'Dib learned rapidly because his first training was in how to learn. And the first lesson of all was the basic trust that he could learn. It's shocking to find how many people do not believe they can learn, and how many more believe learning to be difficult. Muad'Dib knew that every experience carries its lesson."
― Frank Herbert, Dune
Note: These podcasts were made with NotebookLM. AI can make mistakes. Please double-check any critical information.
Ground-Truth-as-Code: Testing AI Agents on Live Data
•Mike Breault
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
0:00
|
5:30
How do you grade an AI agent when the correct answer changes with the data? We explore ground-truth-as-code: an evaluation approach that reruns reference calculations against live systems instead of relying on stale answer sheets. Discover how breaking responses into individual facts helps measure correctness and completeness—and why reliable testing matters for data science agents in production.
Note: This podcast was AI-generated, and sometimes AI can make mistakes. Please double-check any critical information.
You know, uh the other day I was trying to balance my checking account. I had this printed statement from like last week, and I'm going through it line by line, but absolutely nothing is matching up.
SPEAKER_01
Oh, right, because of all the auto payments and stuff.
SPEAKER_00
Exactly. Auto payments had gone through, a deposit cleared, so the live numbers had completely changed, but my piece of paper just sat there, you know, totally frozen in time.
SPEAKER_01
It is a massive headache, honestly, trying to measure a live breathing system against a static snapshot.
SPEAKER_00
And that is the exact headache software developers are facing right now when testing AI agents. So in this deep dive, we're unpacking a new research paper from Adobe called Skill-Based Agentic Evaluation for Real-Time Data Science Tasks.
SPEAKER_01
Right, which sounds like a mouthful, but it's actually fascinating.
SPEAKER_00
It really is. We're looking at the mechanics of how developers are keeping AI agents perfectly accurate when they analyze these massive, constantly updating live databases. It is just a thrilling look at uh human ingenuity.
SPEAKER_01
Yeah. So the core problem the researchers are tackling here is that standard AI benchmarks rely on what we call frozen reference answers.
SPEAKER_00
Okay, so like a fixed answer key.
SPEAKER_01
Exactly. Imagine you ask an AI agent, uh, what were our top sales figures yesterday? The correct answer today is entirely different from the correct answer tomorrow, right?
SPEAKER_00
Right. So if the AI gives you the correct live answer, but the benchmark is frozen on yesterday's data, the AI actually fails the test.
SPEAKER_01
Aaron Powell Which makes the evaluation totally useless. So the Adobe Team S solution is something called ground truth as code.
SPEAKER_00
Ground truth as code.
SPEAKER_01
Right. Instead of handing the evaluator a fixed text answer, the correct answer is actually an executable Python function. That code recomputes the truth from the live data at the exact millisecond the AI is tested.
SPEAKER_00
Okay, let's unpack this. It's like grading a math test. But instead of giving the teacher an answer key with the number 42, you give them the actual algebraic formula.
SPEAKER_01
Yeah, that's a great way to put it.
SPEAKER_00
So they calculate the right answer on the fly, no matter what variables the student was given.
SPEAKER_01
That analogy holds up really well. I mean, the AI and the grading system are finally looking at the exact same live reality. It entirely eliminates the problem of data drift.
SPEAKER_00
That is brilliant. Yeah. And you know, getting agents to interact flawlessly with live environments is complicated work. Which uh brings me to our sponsor. This deep dive is sponsored by Embersilk.
SPEAKER_01
Oh, nice.
SPEAKER_00
Yeah. If you need help with AI training or automation or integration or you know, software development, they are the experts. If you're uncovering where agents can make the most impact for your business or personal life, just check out Embersilk.com for your AI needs.
SPEAKER_01
Awesome.
SPEAKER_00
Okay, so the code solves the shifting answer key. But I feel like this creates a new headache.
SPEAKER_01
How so?
SPEAKER_00
Well, if I ask an AI for a data summary, it might spit out a bulleted list or a raw table or maybe like a three-paragraph essay. How does an automated greeting script actually parse that messy human-like text to verify the math?
SPEAKER_01
Ah, right. They solve that by using what they call format agnostic factoid scoring.
SPEAKER_00
Format agnostic factoid scoring.
SPEAKER_01
Yeah, it's a bit technical, but basically the evaluation script completely ignores the layout. It takes the AI's response and chops it down into tiny atomic claims, which they call factoids.
SPEAKER_00
Okay, making it bite-sized.
SPEAKER_01
Exactly. And then it does the same thing for the code's output. Finally, it compares those two simple lists based on precision and recall.
SPEAKER_00
Wait, I see a flaw here though. If I ask for the sales numbers and the AI gives me the correct numbers, but then randomly hallucinates that our CEO just resigned.
SPEAKER_01
Oh, I see where you're going.
SPEAKER_00
Right. Like is the grading script gonna mark as correct just because the math was right? Does it get penalized for showing off?
SPEAKER_01
Well, the researchers accounted for that. That is where precision and recall come in. Precision acts as a hallucination check.
SPEAKER_00
So it catches the fake CEO news.
SPEAKER_01
Exactly. If the AI introduces extra facts that are wrong, precision drops, and the AI is penalized heavily. Recall, on the other hand, checks for omissions.
SPEAKER_00
Like, did it miss any required facts?
SPEAKER_01
Right. But interestingly, if the AI provides extra facts that are actually true, it's just treated as a bonus.
SPEAKER_00
Aaron Powell Oh, wow. Really? I'm curious about the compute cost of all this checking, though.
SPEAKER_01
It's actually highly efficient. Because the evaluation model is just checking true or false on a simple list of atomic facts. It doesn't have to burn compute trying to read and parse dense, complex formatting.
SPEAKER_00
That makes a lot of sense.
SPEAKER_01
Yeah. When tested on a synthetic production database, this method not only improved agreement with human experts by 29%, it actually reduced token computing costs by 16%.
SPEAKER_00
Aaron Powell That is incredible. We are genuinely solving the hallucination problem here, building AI agents that dynamically adapt to the real world with, you know, incredible accuracy.
SPEAKER_01
Aaron Powell It really does point to an inspiring future for data discovery. I mean, the potential is huge.
SPEAKER_00
It is. And it makes me wonder if we can now encode factual ground truth as live executable code to evaluate data, could we eventually encode complex logical frameworks as executable code?
SPEAKER_01
No, to seamlessly guide AI reasoning in real time.
SPEAKER_00
Exactly. A live compass for logic itself. Now that is something to mull over. Maybe next time I check my bank balance, I won't just need a live calculator. I'll have an AI perfectly calibrated to the live logic of my finances. Truly. Well, if you enjoyed this podcast, please subscribe to the show. Hey, leave us a five-star review if you can. It really does help get the word out. Thanks for tuning in.