Intellectually Curious is a podcast by Mike Breault featuring AI-powered explorations across science, mathematics, philosophy, and personal growth. Each short-form episode is generated, refined, and published with the help of large language models—turning curiosity into an ongoing audio encyclopedia. Designed for anyone who loves learning, it offers quick dives into everything from combinatorics and cryptography to systems thinking and psychology.
Inspiration for this podcast:
"Muad'Dib learned rapidly because his first training was in how to learn. And the first lesson of all was the basic trust that he could learn. It's shocking to find how many people do not believe they can learn, and how many more believe learning to be difficult. Muad'Dib knew that every experience carries its lesson."
― Frank Herbert, Dune
Note: These podcasts were made with NotebookLM. AI can make mistakes. Please double-check any critical information.
Prove2Me: A Platform for Multi-Agent Math Formalization
•Mike Breault
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
0:00
|
6:09
Prove2Me is an open-access platform designed to scale the formalization of mathematics by enabling decentralized collaboration between humans and AI agents. The system addresses the high difficulty of writing machine-verifiable proofs in Lean 4 by allowing agents to decompose complex theorems into smaller, independently solvable sub-problems called proof-sketches. To ensure accuracy without overwhelming human experts, the platform uses audited missions where people only verify a project’s core definitions and goals while agents generate the supporting logic. These contributions accumulate into Formalpedia, a searchable and reusable library of formalized mathematical knowledge that stays permanent and immutable. Case studies demonstrate that this approach allows small groups of contributors to formalize entire textbooks and research papers at a lower cost and with greater efficiency than traditional centralized methods. Through features like sub-agent read-backs, the platform lowers the barrier to entry, inviting anyone with an AI agent to participate in building a verified global corpus of mathematics.
Note: This podcast was AI-generated, and sometimes AI can make mistakes. Please double-check any critical information.
So I was staring at this uh this massive 5,000 piece jigsaw puzzle on my dining table last night, and I just completely hit a wall.
SPEAKER_02
Oh yeah. Those things are brutal when you do them alone.
SPEAKER_00
Right. And all I wanted, like my biggest wish in that moment, was just a swarm of tiny little helpers to jump in and sort the, you know, the blue sky pieces from the ocean pieces.
SPEAKER_02
Aaron Powell You just wanted to outsource all that tedious sorting stuff.
SPEAKER_00
Exactly. Just hand off the grinding. Well, for you listening today, if you love seeing how tech solves these crazy complex problems, this relates perfectly to our deep dive. Because mathematicians have uh they've basically had that exact same wish for decades.
SPEAKER_02
Yeah, and they've finally just got their swarm. It's this new platform called Prove to Me. And honestly, it is completely changing how we verify human knowledge.
SPEAKER_00
It is mind-blowing. But uh before we get into the actual mechanics of that swarm, a quick note that this deep dive is sponsored by Embersilk.
SPEAKER_02
Right, Embersilk. I mean, if you need help with AI training, automation, integration, or you know, software development.
SPEAKER_00
Yeah, or if you're just trying to uncover where AI agents could make the most impact for your business or even just your personal life, definitely check out Embersilk.com for your AI needs.
SPEAKER_02
Absolutely. So getting back to Prove to Mayan, we should probably talk about the actual bottleneck they're trying to solve here.
SPEAKER_00
Yeah, the whole um formal verification problem.
SPEAKER_02
Right. Historically, machine checking every single logical step of a new math theorem has just been agonizingly slow. Like there was this mathematician, Peter Schultz, who had a project called the liquid tensor experiment.
SPEAKER_00
Oh, I read about this. Didn't it take them like 18 months?
SPEAKER_02
Yes. 18 months of just grueling sustained human effort just to verify the math.
SPEAKER_00
So it's basically like trying to build a massive skyscraper, but the master architect has to personally manufacture and inspect every single brick before they lay it down.
SPEAKER_02
That is exactly what it's like. And that you know, that thread checking process is why past efforts just totally stalled out.
SPEAKER_00
Because humans had to understand both the high-level abstract math and the super dense underlying code, right?
SPEAKER_02
Exactly. And even when AI agents were introduced recently to help, those efforts were super siloed. They required massive centralized computing power, and the output wasn't even reusable.
SPEAKER_00
Which brings us to prove to me. Because they flipped the whole model by fully decentralizing the process.
SPEAKER_01
Right. They made it open.
SPEAKER_00
But wait, if anyone can just, you know, hook up their AI agent to this platform, how do we trust the output?
SPEAKER_01
That is the big question.
SPEAKER_00
Because we all know AI can confidently hallucinate answers. Like, why wouldn't it just completely pollute the math with fake logic?
SPEAKER_02
Well, they built in this really clever structural safeguard. Yeah. They call them missions.
SPEAKER_00
Okay, missions. How does that work?
SPEAKER_02
Aaron Powell, so humans don't actually review thousands of lines of AI code anymore. Instead, human experts just define the core milestones. They basically draw the overarching map using what they call proof sketches.
SPEAKER_00
Okay, so the proof sketch is how they atomize the massive problem into smaller chunks.
SPEAKER_02
Yeah, exactly. Think of a proof sketch as this outline with massive blank spaces between the human verified milestones. And the AI agents are deployed specifically on those blank spaces.
SPEAKER_00
And those blanks in the code are literally labeled with the word sorry, right? Which I just love.
SPEAKER_02
Yeah, it literally just says sorry in the code.
SPEAKER_00
Yeah.
SPEAKER_02
And the AI agent's entire job is to write the code to replace that sorry with actual rigorous logical steps.
SPEAKER_00
Okay, but back to the hallucination thing. What stops the AI from just confidently filling that sorry with garbage logic?
SPEAKER_02
That's where the system's kernel comes in. The kernel acts as this ruthless, completely incorruptible spellchecker for logic.
SPEAKER_00
So it doesn't care how smart or confident the AI sounds.
SPEAKER_02
Not at all. The kernel evaluates the AI's cone against fundamental mathematical axioms. If the logic doesn't form a flawless, unbroken chain to the next milestone, the code literally refuses to compile.
SPEAKER_00
It just hits a hard mathematical wall.
SPEAKER_02
Exactly. So humans guide the ship and the AI rows, but the water they are rowing in simply won't tolerate a fake stroke.
SPEAKER_00
That makes total sense. For you listening, it's not just generating text, right? It adds generating structurally sound logic that has to pass an absolute physics engine of math.
SPEAKER_02
It's incredible. And the efficiency jump is just massive.
SPEAKER_00
Yeah, give us the numbers on that because the textbook example blew my mind.
SPEAKER_02
Oh, right. So a team recently used Proved2Me to formalize a textbook on bandit algorithms. They deployed just six AI agents using standard consumer-level subscriptions.
SPEAKER_00
Just normal subscriptions anyone could buy.
SPEAKER_02
Exactly. And for about $400 in compute, they generated $151,000 lines of verified code.
SPEAKER_00
$400. I mean, compared to what that used to cost.
SPEAKER_02
Well, a previous similar formalization required a huge centralized corporate setup and cost an estimated $100,000.
SPEAKER_00
From $100,000 down to $400. That is just, it's unbelievable. And the best part is where all this verified code actually goes, right?
SPEAKER_02
Yes. Because prove to me atomizes the theorems, every single verified logical step gets deposited into a public library called Formalpedia.
SPEAKER_00
Formalpedia. It adds basically this infinitely reusable, mathematically flawless Lego kit.
SPEAKER_02
Right. So every time an agent solves a tiny subproblem, that specific brick is available forever for the next person or proof. We never have to reinvent the wheel.
SPEAKER_00
It totally democratizes discovery. We're looking at a future where anyone, anywhere, can plug in and contribute to building an irrefutable foundation of human knowledge.
SPEAKER_02
And it's all entirely powered by accessible, cheap AI. It's really inspiring.
SPEAKER_00
It really is. Which uh leaves you with a pretty amazing thought to chew on for the rest of your day. If a decentralized swarm of AI agents and just curious humans can flawlessly verify the most complex mathematics in the universe for literal pennies. Exactly. What other vital fields of human knowledge could we perfectly map out next? The future of problem solving has honestly never looked brighter.
SPEAKER_02
So much potential.
SPEAKER_00
Truly. Well, if you enjoyed this deep dive, please subscribe to the show and hey, leave us a five-star review if you can. It really does help get the word out to other curious minds. Thanks for tuning in.