Cut straight from the AMA recording and the team's published talks. Quotes lightly cleaned from auto-transcripts — the audio is the source of truth. Headphones on.
Let us not mention specific names… one of them is hopefully going to be signed this week.
We're not Solana, we're not Ethereum — we're going for the Bitcoin of the AI era.
We are relying on NVIDIA being the core — essentially the ASIC for matmuls.
You download vLLM and it just comes with Pearl mining plugins. That's the dream.
No pre-allocations, no pre-mine. No one — including team members, including investors.
Turning compute into pearls is mining. The endgame is the other way around — pearls into compute.
Money for machines — and humans.
Maybe you can just add structure to the noise.
Vitalik adds: it is probably infeasible. The rest is history.
A proof-of-useful-work blockchain becomes a financing vehicle for AI workloads.
The tapes are the highlights. This is the whole thing. Every public talk and clip we could get our hands on, transcribed in full — sit down and read the record in the team's own words. Auto-transcribed from public talks (so expect the odd rough edge); each one links back to the original. Click any to open.
00:11Hi everyone, thank you for being here Yeah, so let's be frank it took a little bit more than just the math so a lot of the blood sweat and tears of Incredible people in this room Got us here So hi everybody. I'm Omri. I'm the co-founder and CEO of Pearl and
00:33I wanted to deeply thank each and every one of you For being here and for your time and interest in the Pearl project for centuries Money's been a social proxy for human labor and innovation managed through social contracts credit tax and monetary policies Well, it's hard to predict where AI is headed or what are exactly its limitations
01:03One thing that's becoming clearer every day now is that our future economy will be denominated in compute cycles more than it is in in human labor and in this reality where the production of knowledge is Becoming more mechanical in some of the frontier labs in this audience are spending 100 million 130 million dollars a day on inference if I'm correct
01:28One could argue that the real currency is not money It's energy and data the two what we collectively call Compute so the two scarce resources whose fusion creates intelligence is is data and energy data and compute We believe that this dramatic shift in the production of knowledge compels us to rethink the fundamental Properties purpose in the creation process of money what money actually is
01:57The basic scientific question that started This the pearl project and we Elon and I asked ourselves About a year and a half ago was whether we can build a monetary currency Directly from these two key scarce resources data and compute This is the motivation for this question is not merely mathematical or philosophical It is to address at least three core incompatibilities of the fiat system
02:29With emerging AI economy The first the first one is that the unit economics of AI today is simply non-fungible not all LLM tokens are born equal producing a Deep-seek R1 LLM token on NVIDIA hopper has a very different physical cost production cost then producing a reasoning token of chain of thought of Quinn 14 B on AMD MI 3's right so these are just two very different physical costs
03:02The current crude market pricing of you know X dollars per million LLM tokens is at best a crude proxy for that this actual energetic cost But it it hardly serves as as a proxy for the data dimension, which is much more elusive and harder To articulate The second drawback or a limitation is that in the current economy AI consumers are
03:29Completely left out of the upside of the wealth generated by AI so AI users like myself who are you know Just using Claude for for coding or or chatbots. We don't get an anthropic or NVIDIA share Well, most of us at least But by doing so yet consumers are the ones driving demand and model improvement through prompting RL data labeling RLGF and so forth, but we're we don't really have a stake in
04:01The in the wealth generated by the AI revolution, right? So that's that's another part the third one and final one is that in this in this era where Autonomous AI systems are slowly penetrating our you know economy and driving a lot of the value We must have an economic system that allows machines to participate alongside Humans in the same ecosystem
04:28Without you know onboarding KYC and identification, right? That's we just have to allow machines to somehow manage the two only resources we deposit in their hands or in their You know in their in their system, which is compute and data, right? So AI agents just receive two researchers the data the prompts in the in the compute budget And we believe AI agents should be able to transact seamlessly and directly using these resources and to manage them without any permissioned or
05:03Regulatory or social Restrictions with their obviously not subject to Right. So this is what we often call Among ourselves that a money for machines, right? So it's money for machines and humans And we will see how to get all of these three properties and more and then some we will see actually additional Functional properties that were not possible for AI systems that are
05:31In some sense follow from the technology will present in the tool will use to create this AI native currency is proof-of-work blockchains All right, so a Three-minute crash course on on blockchains. What are blockchains? So blockchains are a technology that tries to answer the following challenging question of how do we all collectively?
05:57maintain a Decentralized maintain an update a decentralized peer-to-peer Transaction ledger or network Okay in a way that is consistent that each and every one of us Despite
06:12Having hits its owner her own local view. How can we all maintain and update a consistent state? in a way that's Without any trusted third-party. Okay, so even if 49% of us Are failing or are abstaining from reporting anything? We want the system to remain live right for this to happen each and every one of us should be able to replicate the state
06:40of this data structure of this ledger Right and another feature is that we ideally we would like this system to be permissionless So users can come and go enter or leave the system without any onboarding process So we don't actually know how much compute or how much Users are participating in the system All right, so the first thing you might think is you know, how should we go about it like?
07:05There is no central bank. There is no single point of failure. How should we all agree on the next state, right? This system is useless if we can't if we just maintain it We have to be able to agree on the update rule how do we update and collectively agree on the next state of this chain and The most natural thing you can essentially the only thing reasonable thing to do Assuming you have honest majority which is the weakest thing we can ever hope for right if there's no honest majority
07:33All bets are off. Anyways The most natural thing we can do is just take a cast a majority vote, right? That's the most natural thing to do if there's an honest majority. I will just Well gives a higher chance and then you can maybe repeat it and ensure that an honest The honest majority determines dictates the next State the problem is that you know besides being inefficient. I can't go to each and every one of you and just solicit a majority vote
08:01this rule is actually not secure because by Abstaining from voting some of you who don't think the next transaction is favorable to you Can just collude and abstain from voting that could actually bias the answer to something different, right? So that's actually that's the majority vote is simply not secure The second idea is maybe we can do something more efficient
08:26Which is kind of an unbiased estimator for a majority vote, which is just let's just sample a random leader and let her Or him approve the next state. Okay, so this is kind of an expectation. This is simulating a majority vote But then again, remember, there's no trusted party. There's no single blackboard written in the sky Who draws the randomness? How do we agree on the support of the distribution? Right? It's not remember It's a permissionless system. We don't even know What is the sample space of this distribution, right? So these are, you know, two
09:02You know unsuccessful trials, but actually if you think about it the biggest problem with the majority vote is The fact that if voting is cheap no matter how you implement it if voting is free Then Alice could just replicate herself 10,000 times and just Pretend to be 10,000 voters and she could just bias the result however she wants because you know Zack would communicate with her and he would be led to believe that Alice is acting like 10,000 voters
09:31Does that make sense? And if you think about it, this is you know, there are many ways to view Bitcoin as technology, you know You know, it was a breakthrough in many senses. This is the core innovation of Bitcoin The core conceptual innovation was the guiding principle That making voting expensive is just designating who's allowed to vote and Bitcoin's rule was
09:58voting entails demonstration or investment of a scarce but abundant resource so an expensive but abundant but widely accessible resource Okay, so that was and if you think about it, this was you know, blockchains have evolved in the past 20 years To many other consensus mechanisms, which some of you might know All of them stick to this guiding principle where users or voters or minors will
10:26invest resources Expensive but abundant resources to secure the system in order to incentivize them to invest those resources The system will issue rewards coins or money or however you call it in that will incentivize the the the users to keep securing the system Does that make sense? Bitcoins or Nakamoto's
10:51Observation was that electricity is a resource that satisfies these two properties. It's abundant enough. You know, it's widely accessible in the world But it's definitely not free at least in most countries so Bitcoin's idea was to enforce this voting system or Or, you know Consensus protocol through what's called proof-of-work protocol. So in two minutes, what is proof of work?
11:22So in a blockchain we have Transactions grouped into blocks, right? So in this block we have user ER 20 Transferring 10 coins to user 0 and 5 for example And then each block has a unique ID a unique identifier or a hash code that uniquely identifies it and If Alice wants to add a transaction to this ledger to just approve
11:48She wants to transfer Bob maybe $20 20 coins then What will happen is that Alice would need To update this and the only update operation in the system is the operation of extending a block so You can extend any block in this chain. So it's actually a tree. It's not really a chain But what defines honest users is that honest users Only believe the longest chain the longest path from root to leaf, okay?
12:18That's very it's probably it's going to be intuitive, but it's going to be clear later Okay, so Anything else than believing the longest chain is considered, you know, all bets are off I'm not assuming anything all I'm assuming is that at least 51% of the users are Following the longest chain. All right, and now If Alice wants to add
12:40This transaction to extend let's say Alice is honest So she's trying to extend the the longest chain. So it's like plugging a socket into the To the wall or whatever and she would need to solve a computationally hard problem so in order to add this transaction to the public ledger she would need to
13:02Take the idea of the block. She's trying to plug into to extend She would treat this string as just as a bit string and then Alice's Task in order to approve this block this transaction is to find a completing string this black puzzle piece That will cause a random hash function to evaluate to a very small number Okay, if you know anything about, you know, random hash functions, you would know that any
13:29Guests for the black puzzle piece, which is a bit string has equal chance of evaluating to this number Right and importantly in Bitcoin the number D, which is called the difficulty of the system is dynamically adjusted So that in expectation No matter how much compute there's on the network exactly one Winning puzzle piece arrives every ten minutes. Okay, so it doesn't matter if there are one Hexa flops or 23 tera hash flops in the system the clock the clock keeps
14:06Taking every ten minutes. Okay, so if you think abstractly, what is proof of work? It's just a distributed clock That clicks once every ten minutes every ten minutes. There's exactly One block That's being opened does that make sense? Yeah, and for for them mathematicians in The crowd what characterizes proof of work is it's just a Poisson process where the arrival of the next lottery ticket Is just a Poisson random variable
14:43Why am I telling you all this? So you might think so is the process clear so this is the only operation. This is how the system works Notice that it avoids the problem of soliciting Votes it's completely asynchronous right because I don't care if you know, you don't Avoid voting. It's his loss. The system keeps going right So you might think okay, that's just a random arbitrary
15:11Hard problem random hatching just random number guessing right that's just arbitrary Why not do anything else? But if you think about of it? There are two key properties There are crucial about random hashing which makes it work the first one is that the proof of work puzzle this puzzle piece that Alice needs to solve that miners need to compute it better be hard to compute But very cheap to verify right so if Elon finds this puzzle piece You know finding it is exhaustive search, right? It's just a random number guessing just every ticket is has the same winning chance
15:46But if Elon just shouts out this number I immediately know it's right Yep, the second thing is that in order for the system to be open truly open and fair There cannot be easy instances for Alice right there cannot be like low-hanging fruits because then a dishonest miner Will just pick the easy instance and it will mine it will just scoop rewards much faster And then we just lose control over the clicking of the clock right every economy where you cannot control the pace of
16:20Minting new coins will collapse right so that that that won't work. Is this clear? so these are two very crucial properties and Sorry, so yes a more mathematically. I would say that any miner or user that Spent holds in his hands P fraction of the total hash rate or the total compute in the network
16:46Should have a chance of winning which is essentially P So if you have P fraction of the compute it's only fair It's the only fair thing right if you have a P Fraction of the compute in the network You're winning you're gonna scoop a P fraction of the rewards right and that makes us completely fair in Symmetric, so the the hardest is just uniform for everyone
17:08Okay, and turns out random hashing satisfies this and You know it's been what like almost 20 years Bitcoin is past the test of time whether you know if you ignore The bumpy road of the price it it is a pretty substantial human achievement that such a system actually It actually works right it works and it's slowly getting gaining more and more Approval higher up But the obvious problem is that Bitcoin is just
17:40Consuming as you saw in the in the video a ton of energy just on random hashing right so between 1 and 2 percent of the world Global electricity more than I think Argentina and Norway is spent just on these random number guessing Right, which of course are securing the chain, but beyond that it's just garbage, right? So there's no there's no extra additional utility and if you think about it this problem is even worse today because Bitcoin and AI are actually competing over the same electricity their mind
18:11They're being operational by very different hardware, but they are competing at the core of it for the same energy Right, which is why a lot of the hyperscalers Are replacing today their Bitcoin racks with GPUs and TPUs because the profit per gigawatt is higher in AI today as it is This is commonly referred as Bitcoin's declining security budget where aside from AI as the chain is Becomes more mature and less and less reward are issued
18:41The incentive to keep mining and participating decreases So which that brings us to the you know extremely natural idea Which we were far from the first to conceive of proofs of useful work, which is that you know Most natural idea is why can't we replace random hashing with something with a tacit actually useful, right? that with electricity that's already being spent to do I know drug discovery or Quantum simulation or scientific computing and
19:11This challenge was acknowledged by many thought leaders in the past 10 or 15 years In fact Elon's advisor in the the Weitzman Institute was the first to acknowledge this problem the very first proof of work paper way before Bitcoin one of the So there's been a long line of work on this each and every proposal for useful proof of work was shown to fall short of
19:37satisfying one of those core Definitions or properties of useful work and indeed none of those attempts have succeeded One of the most vocal votes was Voices was by Vitalik the co-founder of Ethereum who said that if we can find some useful Computation which is easy to verify then crypto mining Could be could actually become a huge boon to society not only removing
20:05Subjection that bitcoins wastes energy, but also being societally Beneficial and then a few lines later Vitalik ads this problem is is probably infeasible in a few months later A few years later a theorem switches from proof of work the proof of stake and the rest is history all right, so this brings us to the
20:27the pearl kind of Area where What we're trying to do is essentially a following so consider any real-world problem a protein folding DNA sequencing quantum simulations Vector databases anything you're you're working on and it's being deployed in the real world regardless of blockchains and Suppose this task as f of x you're trying to compute has worst-case complexity t
20:55So worst-case complexity is just the running time of the algorithm that solves any instance of this problem So for sorting n numbers He is roughly n log n For multiplying two n-bit integers He is roughly n squared If you went to Stanford, you probably know that you can multiply numbers faster using FFT, but we'll get into this
21:17So this your n squared is roughly The best Practically like the baseline will use here and what we're trying to do here is We're trying to piggyback on the same Native algorithm or the native computation that computes f of x that's gonna happen regardless of the blockchains, right? People are gonna keep multiplying numbers and sorting things regardless of anything
21:44We would like to have with very little extra overhead Beyond the worst-case complexity. We would like to have two outputs. So two for one Right, we would like to have to produce the output the output f of x which the algorithm does anyway And it produces a lottery ticket a certificate exactly like Bitcoin which has this Poisson behavior Does that make sense? great and
22:11If you think about it The reason proof of useful work was so challenging was unsolved for the past 20 years is That unlike random hashing Real-world problems just they simply don't have the structure that or the verifiable structure that random hashing does right? For the system to be truly useful. I better allow the miner itself to pick its own her own favorite instance x right the input That's very different than Bitcoin Bitcoin you get the assigned the problem by the network and proof of useful work
22:46You're truly useful Proof-of-work chain the miner itself is picking the input right so Alice can choose x at her own discretion And if you have some external gains from it like training inference quantum simulations and good for her She'll piggyback and just on those rewards, but it's crucial to allow Alice to choose her own input Of course, you know real-world problems just don't have this uniform hardest They vary a lot in their difficulty let alone for dishonest players who don't care about the useful word and they can just fabricate
23:20Inputs right so in the multiplication example Alice could just choose a and b to be the zeroes right multiplying Zero by zero is not doesn't take n squared time. It's just trivial Does that make sense so this lack of uniformity? Means the system is going to be very gameable and that's the core challenge of proof-of-use for work Does that make sense? so in the in the paper that you line and I published
23:50About a year ago what we studied is the is the problem of Implemented proof in Nakamoto's consensus on top of a specific task, which is matrix multiplication Why matrix multiplication? I guess I was at the time in Nvidia, but probably The reason was was deeper than that and you'll see in a moment So Matt Maul is just Under the underlying technology of not just training an inference an RL and post-training anything you can imagine
24:20It's probably the underlying Operation probably between 70 and 90 percent of any industry scale operation you can imagine You know self-driving cars robots compression quantum simulations navigation everything in the world is is Matt Maul is the most optimized operation in the In the the history of humanity
24:44So what is matrix multiplication, I'm not gonna get too deep here basically we have two arrays of n by n numbers to simplify Right and we're required to output another array Which contains the product of every row in the left matrix a with any column in the right matrix B Okay, so we have n choices for the rows and n choices for the columns of B Each product has n numbers so it requires n time if you if you multiply all these you get n cubed complexity So this is what's called naive matrix multiplication. That's the that's the algorithm that in video AMD cerebrus edge everyone uses today
25:25There are theoretical faster algorithms, which is part of my academic research, but that's that's not That's not Practical anywhere soon and even if it will be pro will be able to adapt to this but this is something we can take offline so the way GPUs multiply matrices right in
25:46in Normal LLMs the matrices could have like 40k by 40k dimensions, right? And so the GPU would actually tile the matrices So you will split into blocks like r by r tiles It would just load the tiles and multiply talk. That's that atomic unit of operation okay, so you can do those things just by
26:07Multiplying instead of a single an entire row in a column you would just multiply a tile by a tile Which is a smaller mathematical right and that fits into memory and that that you can do All right, so in this paper that we we wrote We showed how to generate a mathematical proof which it has exactly mathematically the same Distribution of Nakamoto consensus but
26:33With a tiny Overhead beyond the worst-case complexity of math mode you can actually Produce those the same security mechanism of Bitcoin What's the math behind it again? It's just two minutes and I'll I'll try to take it easy on you so
26:52What's the naive attempt to do this right so now instead of there's this added dimension of data that Alice gets to choose She picks Alice gets to pick a and b if Alex is Alice is running I know Claude or you know or or or Quen or Kimi K2 she would probably choose a to be the weight matrix of some layer and B to be the activation matrix, okay? and
27:20But she can choose anything she wants right remember no one's you know, it's at completely our own discretion Alice could choose Matrices she pre-computed at home a week ago. No one will know right. There's it's a permissionless system. There's no way to know and The most natural idea is to just do the following instead of hashing in order to extend the block as we saw a few slides ago Instead of hashing the idea of the previous block Why don't we just hash the result of some useful computation of the mathematical right so we just as will compute the mathematical and
27:58I'll just apply the hash to the result of this useful computation. That's negligible compared to the cost of mathematical The problem with this is again if a and b are Unpredictable or random let's say then this actually works works perfectly So if a and b are like ran a hard matrices to multiply this would work the problem of course is that again Alice is free to choose her favorite a and b so she could choose the all-zeros or diagonal matrices or Whatever she wants and the problem is that this will allow Alice to reverse engineer
28:33The the winning criteria because everything is deterministic right there's no randomness in the system So Alice could just reverse engineer everything that you could just keep scooping rewards very fast and game the system by choosing easy Or predictable and I hope this makes sense So the idea is that we must somehow bake in or embed randomness into the native algorithm that computes math mode Because if there's nothing unpredictable in the math mode algorithm Alice can just pre-compute it at home and just game the system does that make sense and
29:10The natural way to do this is to just add a little bit of noise to a and b right that's a very natural thing to do and Indeed This is what we would okay So this is what we're trying to seek here for is what theoreticians call a worst-case to average case Reduction which has a very succinct overhead parameter So we would like to take any a and b that Alice produces and we'd like to compute to reduce it to the task of
29:37Random a times B without paying much more overhead And the first trial is to just maybe we just add Random noise matrices so now no matter what Alice chose even if she's trying to game the system and choose the all-zero matrix as a and B If we add random noise matrices Alice is doomed. That's gonna take n cubed Operations to multiply because random matrix and random Gaussian not moles is n cubed hard. We can take this offline
30:11So now we do the proof of work on the perturbed matrices that's hard right so great That's that has bitcoins like security thing But of course the what's the problem the problem is if Alice is an honest user running AI models She doesn't care about the perturbed product. She wants the original a times B the product of the white matrices So what should Alice do in order to decode the white matrices? She should just peel off the really just open the brackets
30:39She just needs to peel off three extra math moles right right to get the original result She would need to do extra post-processing operation. Just so what did what did we do here? Like Alice paid four math moles to get just one so the overhead here is like 400 percent. It's like x4 We're looking for Little of one overhead like very very small right no one would use this this system No one would pay like four math moles. You can actually squeeze. It's a nice exercise. You can squeeze out
31:08Well, it's definitely gonna take at least two or three math moles. So that doesn't work. Okay so The real Protocol the eventual protocol which we call pearl gem. I think this is the probably pearls biggest innovation Is to follow the same idea, but in order to make the peeling of the noise the decoding cheap Maybe you can just add structure to the noise. So instead of just using
31:39brute force random Gaussian matrices What we can do is we can just add the noise matrices which have They look like kind of kind of margin marginally random, but they actually have structures So if we add what's called low rank random Gaussian matrices low rank matrices are still end-by-end matrices if you open them up But they have linear correlations. So actually you can decompose them into the outer product of two like
32:10You know strip matrices and The thing you should know is that this completely solved the peeling problem because now if we ask out if now Alice Needs to peel off those extra three products if you do the math You'll see that all those peeling products are gonna be low rank rank our matrix multiplication That's gonna cost our n squared which is negligible compared to our cube. Okay, so in practice Elon and and Vipple will talk about this. We'll see what's the actual overhead in practice. That's great because now peeling is just negligible, right?
32:48But I claim this this screws up another thing This protocol is no longer secure because if Alice now chooses the all-zeros matrix matrices Alice is left with a much easier problem than an honest player that's multiplying dense matrices Alice now Just needs to multiply low-ranked matrices rank our matrices Alice can do this operation in our n squared time Whereas an honest minor running actual LLM math moles is gonna pay n cubed
33:20That's completely breaks the security. The system is no longer fair, right? There are easy instances. Does that make sense? It looks like we're trying to hold the string from both ends We tried to make the system secure So we added we added 300 percent overhead and we tried to reduce the overhead We screwed up the security, right? It seems like fundamentally this you know this trade-off seems inherent And I think pearls, you know the key innovation here the key conceptual idea was the following that
33:50It's you know for an honest minor for an honest honest Alice that Wants to run LLMs all Alice cares about is the output is a product of a and b right? She doesn't care about the intermediate computations But no one said we need to do the proof-of-work on the output we can we're free to do this on Intermediate computations as well, and I think this is the key idea. So we would like to prevent Alice From exploiting the low-rank structure of the noise in the intermediate space, right?
34:19And this is how we do it so instead of just verifying or hashing the output of A time of a prime times B prime which is gameable. We just saw this is easy Instead we're going to verify we're going to do the proof-of-work Not on the output, but on the intermediate computations that the GPU already does so when the GPU multiplies those to our by our tiles We will just for every such our by our tile that the GPU is going to multiply anyway. We're gonna
34:51Force Alice to hash the product the the result of those intermediate tiles and Check if it starts with let's say 19-0. This is exactly bitcoins adjustable difficulty have a puzzle problem So so and the nice thing is that about low-rank random matrices is That marginally every our by our tiles that you see here is going to be uniformly random This is just a our wise independent distribution. So that's a standard property of course, there's tons of correlations between those styles, but
35:27That's that's the hardest conjecture and then that that's the protocol. That's it. So It was a little quick, but that's it. So Alice just submit commits to their matrices No matter what she does We're gonna add low-rank noise and then we're gonna verify it by having every GPU cycle that multiplies those Intermediate tiles add an extra hashing operation and check so think of it at every our by our tiles Gives you an independent lottery ticket to open a block
35:57This is what Elon is going to talk about later in practice And the hardest assumption is that indeed in order to open all those blocks you need and Cube time which is the naive thing. So we got both peeling which is very cheap and we got security assuming this this Assumption here which has been tested by a lot of world-class
36:22Mathematicians some of which are in this crowd All right, so just to finish up. I have like three minutes What we did in the past year is take this protocol from white paper from the math all the way to production What it means that we actually spent most of not we like most of the pearl team here spent most of his time in the past year just Coding up those kernel those kernels and CUDA and soon AMD and you know
36:50we're slowly expanding Elon will say more about this and what we have today is a is a software package on GitHub which probably some of you have seen which a lot which is you know, which has a VLM plug-in and allows you to run those two-for-one LLM Modules and Really the implication of this system
37:15Which is you know, we're slowing trying to understand this this will be more apparent in Raphael's talk Later this this morning, but the obvious application of pro is just funding the AI build that right So they're reducing the price of training and inference because now using those two-for-one CUDA kernels every GPU cycle that you know
37:42Every new lab is gonna or hyperscalers gonna run can with a very small overhead Which you learn will discuss soon Can just do two-for-one so you have a parallel revenue stream for every GPU cycle on earth It's very important to realize this is not a verticalized system, right? It's not only for training or only for inference or only for closed source or only for open source data centers edge Transformers diffusion models world model doesn't matter everything is not bulls. So the I think the strength of pearl is that
38:15The fact it's not verticalized. It's just an index on all an AI and all of its forms, right as long as Matt most is still around Pearl should be part of that the total addressable market of pearl is According to Eric Smith is like 90% of the electricity in 10 years. So it's I think that's the The strength of pearl and if you think about it this thing, you know after we crack the technology It creates a completely new economic model
38:46Which people just economists haven't thought about because now you have this proof of work incentive but you have also external incentives from being paid for training and inference in the equilibrium dynamics of this system and Implications the market cap of inference and training in AI are completely new things some deep Mathematical questions, which we'll hear from in Rafael's talk in an hour or so or like half an hour And that's it. I think
39:16you know You know Presumably the most native miners or users of the pro network are going to be autonomous AI systems and If you think about it, this is going to be the only currency that AI agents insist in AI systems can natively produce without any KYC any on-boarding process any Regulation is they just take their computing data budget. They just manage it completely seamlessly
39:42So that's I think a unique feature of the system Of course, it has more unique applications that I won't get into like essentially what pearl is doing It's implementing a decentralized voting system for AI for all of AI for all of AI agents systems in the world Because you know if you just go beyond transaction ledgers and decentralized banks what we did here is a completely generic way to Implement a majority vote. So that's that's something that we're also exploring. That's it. Thank you so much
00:05Hi, everyone, my name is Rafael Paz, and I will be talking about the economics of proof use for work. So this started some time ago when Elon and Omri came to me and told me about this amazing protocol, and I loved it, and they said they've been talking to a lot of bunch of people about it, and everybody that hears it loves it, but they often get the complaint or concern that if there is a useful proof of work, and you build a cryptocurrency around it, then the value
00:41of the currency should be zero, because as everybody knows, the value of a proof of work currency should be proportional to the effort and the cost it takes to mine it, but if I could also reuse the work for something else, if the work is useful, then I can sell that work, and that effectively makes the cost of mining zero, and therefore the value of the currency zero. So that seems like a pretty bizarre paradox that the better the technology becomes, the less
01:14the currency should be worth. That's sometimes called the proof of use work dilemma, and it seems like a pretty fundamental paradox. I like paradoxes, so this intrigued me a lot, and as with all the best paradoxes, I hope to convince you that once you look at it the right way, the paradox disappears, so that's what I will talk about, but before telling you how the paradox disappears, let's just briefly
01:43recall the problem we're trying to solve, that of building a blockchain, a secure blockchain. So as you heard from Milan and from Omri, a blockchain is just a way to implement a trusted public ledger, and this ledger should be robust to a single point of failure. So how do we achieve something that's robust to a single point of failure? Well, there is a standard approach, if you want to be robust, just replicate all your data, right, trivially just replicate your data, a bunch of different servers, and then
02:15if someone can go down, we still have the data. Of course, once some of them go down or some of them broken into, we don't really know which is the right one anymore, that's always the problem with replication, right, how do we find out which of the replicas are actually correct, and of course, there are standard methods we just vote, right, and we say that if a majority of them are still fine, then the majority vote will be correct, and we can recover the right result, right, very easy.
02:47So we'd like to have a system that is robust to faults and attacks as long as 50% of the replicas are okay. Now as Omri beautifully explained, there is a problem with voting. It seems trivially, you know, correct, but if you want to have an open system, a fully permissionless system, then an attacker can just spawn lots of different nodes, and votes are really not a very reliable approach, right, and that's why in normal election,
03:18once we take off everybody, right, when you come and vote, we take off your name, we check your ID to make sure people are only voting once, but in a permissionless system, we can't do that, and that's where Nakamoto's brilliant idea behind Bitcoin came in, and Nakamoto said, look, let's not count votes from people, let's count votes from computations, right, so each time you spend some computation, you get a vote, and this is what's referred to as a proof of work.
03:48So you're expanding some computation, every computation gives you a vote, and now if we assume that 50% of the computation in the world is being honest and not broken into, we can actually implement this majority of voting in a correct way, okay. Now this proof of work is often called mining because now people are in order to vote, need to mine, do this computation in order to get votes, okay, and this mining is implemented by, as we heard, just having people trying to solve useless puzzles, so they're trying
04:23to find useless puzzles, and once you've found a solution to a puzzle, that's going to be your proof of work, you're proving to the world that you have expanded a certain amount of computation, okay. Brilliant idea, and in such a scenario, we can now get a notion of security, which hasn't been said to break the system, you need to control more than 50% of the computational power in the network, okay.
04:52This intuitive sense correct, it can also be proven, it's actually quite non-trivial, and in a result from roughly 10 years ago, we managed to prove that, indeed, Nakamoto's protocol is secure as long as the attacker controls less than 50% of the computational effort, okay. Great. Now, you might get to the question, why is it the case that people actually doing this computation?
05:17There's a lot of computation being spent to make the network secure, but why would the honest people want to do it, the attacker maybe wants to do it in order to break the security, but why do good people spend all this computation, and that comes from the fact that this brilliant protocol is designed so that every time that you have expanded a certain amount of computation, you actually get some block reward, so you're getting paid for doing the computation, okay. And that's really the key economic insight behind, like modus billion protocol.
05:49So you're getting some bitcoins, every time you solve this proof of work, okay, today we get three bitcoins roughly, and that's called the block reward, and I'm going to refer to that as all. R is three, in parallel R is around 2.7 thousand at the moment, okay. And the block reward is how many bitcoins you get, and then if you want to translate into the dollars, you need to multiply by the price of the, of the Bitcoin, right.
06:15So the block reward in dollars is P times R, and now for this to be meaningful, for honest people to want to actually run this protocol, we need to make sure that the block reward P times R should be at least as high as the cost of mining, otherwise nobody wants to do it, right, because mine cost a lot of electricity in the GPUs, so I need to be able to get rewards that are higher than that, make sense? Okay, so these two things together give you a very nice notion of economic security.
06:49In order to break the protocol, you need to control more than 50% of the computation spent to keep it secure, okay, and that in particular means that in order to break the amount of money you need to spend needs to be at least the block reward divided by 2, because with the block reward needs to be at least as much as the cost of keeping it, of running those computers. All right, to summarize, in Bitcoin, mining requires massive amounts of useless work, we're
07:21just solving these crazy puzzles, okay. But the intention actually is that the uselessness of this work is that makes the system secure. To break the system, you need to spend 50% of that computation, okay, and that in turn needs to be very expensive, okay, so you need to break, you need to wait a lot of resources. Okay, enters proof for useful work. So can we create proofs of work that are not just being completely useless, but are actually
07:55doing something meaningful? The pro break-through that you heard Omri and Ilan talk about shows that actually surprisingly it is possible, and it is possible to actually get proofs of work for not just crazy hashing, but for matrix multiplication, the very computation that really fuels the whole evolution. So now there's no need for wasted computation, we just take those proofs of useful work, plug it into Nakamoto's blockchain protocol, and we get a secure blockchain.
08:27Without wasting computation, awesome, but does it work? In fact, if you go back to the previous slide and recall why did I say that this was secure, the argument was that to break security, you need to waste a lot of resources. But if the resources actually are free because they're being used for something else, then it might not be costly, and security is gone, right, and indeed this is a fault-clear argument, it's been used in the literature.
09:00That says that proof-of-use work is fantastic as a conceptual, philosophical idea, but maybe it's not the right tool for security and blockchain. So let me go over this argument and to just spoil it already, we'll see that this argument actually is flawed, okay? But let me go over it first and we'll see if you can find the mistake in it. So recall, in Bitcoin, in order to break security, this attacker needs to spend a lot
09:27of computation, more than 50% of the computation is secure in the network. And recall, that computation needs to be expensive, it needs to be at least the block reward divided by two, that's the cost of it, or rent in those computers, okay? So to break security, you need to spend a lot of money. Now, let's consider a scenario where we have a proof-of-use work instead. This attacker still needs to rent all his GPUs in order to break security, right, it needs
09:52to still control 50% of the network. But attacker's not just breaking the network, he does all his useful work, he goes out to the market and sells it, right, this useful work can now be done to, can now also provide AI inference and he can go and just sell it to the market. What's the price he gets? Well, he was going to get the price that is like the current market price for compute
10:18and that's exactly what he's spending to rent those GPUs. In total, what's the cost of this attack? Anyone? Zero, right, he runs computers and he uses those computers in order to provide AI workloads. So at least if you have a proof-of-use work with no overhead, right, the optimal proof-use work, it gives you a fully insecure system, pretty convincing, right?
10:49So we'll see today that this argument actually is flawed, okay? So to explain why, let me start by an analogy and then we'll jump into why this analogy applies in this context also, okay? And this analogy is going to be, we're going in plain chess. So let's say I'm going to go and play chess against a chess master. I'm okay at chess, but I can't compete with a chess master.
11:15So I go and play one guy in the morning, yeah, I lose, okay? I feel emboldened and I decide let me go and play another guy in the evening, I will lose again, okay? It doesn't work. Next day, I decide I'm going to play them at the same time, okay? I'm not going to play them isolation.
11:38I'm going to let these two games be happening at the same time. That seems to make it harder for me to win, but actually not. When this guy makes a move, I go to this game and I do the same move. Then when this guy makes a move, I go here and I make the same move. And I'm affecting making these two guys play against one another. And what happens, I win at least one of them or get a draw in both.
12:06So suddenly, I managed to do something that was impossible in isolation, right? So what's the lesson learned here? The lesson learned is that actually we cannot be treating economic systems in isolation. That's what the flawed strawman argument did. In fact, when you have two economic systems and you have some action that can operate them both at the same time, which is indeed what this two for one action at polls has created
12:36is the way. That creates some interactions between both these markets and that leads to amazing things to happen. We cannot just think of them separately. And indeed, let's look at what happens once you think about these two markets, not separately, but allowing some interaction between them.
12:58Let's go back to this attacker. We said he's spending all his computation, he's generating a computer to break the network, and he's going to sell it. Now this analysis thinks about the two markets being completed, he's doing his stuff here, and then there's an inference market that is pricing inference, but that's wrong. If he's making use of it, then so should the good guys producing inference be doing.
13:30Inference providers should also be using a two for one operation, right? Just as he's doing it, so should they. So they should actually be mining in an efficient economy, and they are getting now pearls to subsidize their inference. What that means is that, in fact, the price of inference should decrease. How much will it decrease by?
13:55It will decrease by exactly the block reward in equilibrium. So what happens here is that when he's going out there to sell the useful work that he created, he's not going to get the full production cost of that inference, he's going to sell it cheaper, he's going to sell it for under-production cost, because these guys are actually receiving pearls that are bringing down the price of inference. So in order for him to actually do this attack and sell it, he can only sell it for the discounted
14:31price, so he still needs to pay the rebate. And it turns out that if you do the math, the rebate is going to be exactly the block reward. So in summary, in fact, the price of attack turns out to be exactly the same as in Bitcoin. It's going to be the block reward divided by two, exactly the same numbers in Bitcoin. It turns out that actually, as you'll see in the full economic model, the block reward is going to be at least as high as that in Bitcoin, and sometimes even higher, so the cost
15:01of attack will actually be even higher, in some cases, in Bitcoin, but never lower. So to summarize, AI providers are now getting tokens whenever they provide AI compute. Those tokens enable them to now undercut the current price, and in a free market, that means that the price of inference is going to go down. And that means that an attacker now needs to still pay this discount on inference to attack the system, and that discount is going to be exactly the same as, in fact, the block reward
15:32before. All right, so to forenose all of this, we need to have an economic model of proof of user to work, and that's something that we've been working on. And in roughly a week or so, we will release a paper on this. And in essence, the very formalized thing is by considering two different markets, a security market, and this is a market for, a mining market for maintaining a blockchain.
15:58And a second one, an AI market, or an inference market, because we have these two markets. And then we consider a bunch of players that are just compute operators, GPUs. A player can choose between three actions. They can choose to just mine, to just provide their compute in the, in the blockchain. This is just like going in mining Bitcoin today. Or they can do the max in L for learn.
16:27They can just provide user GPU for product inference. That's a classic world, or they can do this thing referred to as a duplex operation. The duplex operation is a joint, a two for one operation. It participates in both of these two markets at the same time. It co-producers both mining and inference, but maybe with some overhead. So it doesn't get as much as if you just do this, and they're separately, but you can
16:55do them at the same time. And the question is what happens in such a scenario when considered these two joint markets with this action that connects them. And the main results of the paper are going to be, number one, the cost of attack is actually at least as high as in a Bitcoin style world that we refer to as Bitcoinia. And the main, second main result is that we can actually get a complete characterization
17:20of the equilibria in such a scenario as a function of the duplex overhead. So without going to the formulas, let me just kind of highlight what happens in this scenario. In a scenario where you have very good low overhead, what the economic effects of a two for one or duplex operation that enables doing both of these two mining and inference at the same time with low overhead. In that scenario, the blocker words that are being created in a proof-of-work blockchain
17:56are subsidizing AI workloads, and that is always going to lead to lower prices for inference. Now under so-called elastic demand for inference, or something called the Jevons effect, that will actually lead to not just more inference, but even a larger inference market. Let me just mention a word about the so-called Jevons effect. Many of you have probably heard about it. The Jevons effect, in essence, says that once you improve the efficiency of a resource,
18:31that leads to not just, it actually leads to a larger total consumption of the resource. So this is something Jevons already noticed in 1865, that once steam engines became more efficient, people thought, now the price of oil is going to go down. And instead, the opposite happened. There were more need for oil prices going up. Why would that happen?
18:55Well, the conclusion was that once engines became more efficient, suddenly that unlocked new use cases for steam engines, and in fact, we needed even more oil. And the same is expected to be true for the AI market. Once the price of inference becomes lower, that will actually make the size of the market grow because it unlocks new use cases. Okay.
19:19So let's jump in for a few minutes. Let's spend roughly five minutes and start to give you some details of the model. So excuse me a little bit for, I'm going to give you some equations here on the board, but hopefully, you'll be with me for a few minutes. To formalize that the full model, let's first study a pre-poil world. A world where there is no two for one yet, no duplex operations.
19:44So what does the world like that look like? Okay. A bit-constant market where people can do what are referred to as solo mining, they're just doing mining, okay? And that is going to produce some crypto tokens. When I say tokens today, that refers to crypto tokens.
20:01And let us the note of the number of people that are actually doing that. And then we have a separate machine learning market, an AI market where people are playing this action L, and that is producing some useful work, okay, just inference or AI. These operations both have a cost, it's the same cost E, E refers to the cost of running a GPU for 10 minutes, okay? So we're thinking of block times of 10 minutes, E is the cost of running a GPU for 10 minutes.
20:31What are these markets? What do these markets look like? So this is now pre-poil, this is, you know, we have a Bitcoin market and we have an inference market. These are well-established and well-studied markets in the ecosystem. Let's look at the Bitcoin market, okay?
20:45The Bitcoin market is something called a Talok market or a cake cutting market, okay? In Bitcoin, we're getting a block reward every 10 minutes, so we're getting these P times R dollars, let's say, whatever it was, $3 times 80,000, $20,000, $40,000, 10 minutes is being created, okay? And we're doing a cake cutting contest. So that means people can enter, any GPUs in the world can come and enter.
21:12And the more people are entering, the smaller part of the pie I'm getting. So that's what this picture shows here. So the reward I'm getting by entering this is going to be P times R, the block reward, divided by S, the number of GPUs we're entering, minus the cost for me to run a GPU, okay? And now we might ask, what's going to happen in such a market? What's going to happen is that if this number is positive, then more people are going to enter,
21:40right? If I'm making more money than the cost of producing it, more people are entering, and that is eventually driving up the number of people here to become exactly so that P R over S equals E. And that's what's going to happen in the equilibrium, great. So very easy way to solve what happens in a Bitcoin world. Everyone with me?
22:02Yeah, great. Now we move on to the inference market. This was even easier. That's something called a classic betron market, okay? This may be one of the most classic markets in the literature. So we have a demand curve that specifies for a certain price, what is the amount of compute
22:23that the market demands if the compute is being sold at a certain price P, okay? And it's very easy to see here that in such a market, the price eventually is going to converge down to exactly E, the cost of producing it. Because if I'm selling it for $10, but the actual production cost is $5, then it will be another guy that will sell it for $9. And then I will sell for $8, and we go down until we're actually going to end up at exactly
22:53the price of production, okay? So the reward that GPU is getting is the price we're selling it for minus E, the cost of producing it. And in equilibrium, P is going to be equal to E. And in that case, I know exactly how much inference has been produced, that's just this demand curve of E. Very easy, okay? This is the people world, okay?
23:17I want you to think of this mining operation as a fork. This fork is, I'm using it to dig into the ground, to dig for some crypto tokens, okay? And I want you to think of learning as a spoon that allows me to scoop inference workloads, okay? And now enters this duplex operation, this two for one. And that's like a spork, okay?
23:40This allows me to dig in the ground for coins, and allows me to scoop AI workloads also. But it's not as efficient during either this task. There is some overhead here, okay? I'm going to refer to this overhead as alpha and gamma. So alpha is the overhead with respect to how much worse duplexes are doing machine learning. And gamma is how much worse it is at mining.
24:02Right? And in the construction that Omre and Ilan showed, notice that you had to add some, in order to do machine learning loads, you need to add some extra operations. That you could strip away a little bit of overhead, okay? And the same way, if you only wanted to do mining, you could have started off with zero old zero matrices, and that would make the computation a little bit easier.
24:23So there are some overheads involved in doing this, okay? And that's what alpha and gamma are correspond to. So mathematically, using this duplex operation, it's going to correspond to doing a one over alpha of a learning operation, and a one over gamma of a mining operation, okay? So I'm getting that as same output as doing that, but at the price of only one compute. So for example, if you think of alpha and gamma, these are just examples to make the numbers
24:57easy. If you think of them as being 1.3, so you have 30% overhead. That means that doing this duplex operation is going to give you 80% of what you want to do just pure machine learning, and 80% of doing pure mining, okay? So you're effectively getting 1.6 for one, makes sense? You refer to as a proof of use of work as being non-trivial, if the sum of these things
25:24is more than one. You're getting more than one for one, all right? You're getting even 1.1 would be 1.1 for one, that's only again, you're doing two more than one at the same time, okay? And now we can set up these reward equations taking into account these duplex. So in this, the fork here, the mining operation, as I said, the reward I'm getting from mining
25:48is just PR over S minus E, that's from before the reward of selling inferences, the price of the inference minus E, okay, as before. And now we have duplex, and duplex is going to, whenever you use a duplex, I'm going to get some of the inference that I'm producing, so I'm going to get the price, but divided by alpha because there's overhead. And then I'm going to get some of this thing, the PR over S, but divided by gamma, okay, and
26:17then I only have to pay E, okay? So now I can just set up, what are the prices, what are the rewards that people are getting? And then in equilibrium, as we said, nobody should be able to make any money because if people are able to make money, then more people are entering to drive down the prices. In equilibrium, none of these things should be more than zero, and furthermore, if some action is active, if I'm actually doing something, I'm not making negative profit, nobody wants
26:44to do something that you lose money. So those are the equilibrium conditions, these are classic equilibrium conditions. And other questions, we have these conditions, let's just solve what equilibrium is. And that's typically a non-trivial thing, because now, you know, typically people analyze things by saying let's consider them separately, but we can't do that now, I mean, that was the flaw for the previous argument that people said like, you know, let's assume that this
27:08mark is operating independently, it's not, we need to actually, in fact, consider the con joint production, that we're producing things in two markets at the same time. So the main result that we show in the paper is that there exists, in fact, a perfect characterization of equilibrium. There is a one-dimensional characterization, and there's a single parameter that we refer to as the TI ratio, token inference ratio, and here token means crypto token.
27:36This theta, think of it as the, or think of it, it is, it's the ratio between the size of the cryptocurrency market versus the size of the inverse market. Once you tell me the ratio between these two sizes, that perfectly determines the equilibrium, as a function of the, overheads of the, of the primitive. So without going into the math, let me show through a, through a graph, what happens. In fact, we get a few different interesting phase transitions.
28:10So here's a graph that shows how compute changes as a function of the efficiency or the overhead of the duplex operation, the two-front operation. So here we're starting with something that has low efficiency. In that case, we're in a world where we refer to as bitconia. In bitconia, that's the pre-poly world. People are doing pure mining, here, this is wasted work, they're just doing mining.
28:39And here people are doing inference, okay, everything is completely separate. Now once the efficiency of the two for one becomes, so nobody's using this two for one. As if it wasn't there. Now once the efficiency becomes better, we enter fortasia. In fortasia, nobody's doing any, any pure mining anymore, pure mining disappeared. And instead of the impure mining, people using the pearl two for one operation instead, starting
29:08to use duplex. So duplex still has overhead, it's not a no overhead here. And here we see that in fortasia, this light blue area here corresponds to the amount of compute that is actually not useful, because of the overhead. So the overhead here is actually the same size as a bitcon. And indeed, the price of the coin is going to be the same.
29:35And the amount of inference delivered is actually also the same. So in terms of an economic market, it's not changing much. The price is the same as a bitcon, the amount of inference produced in the world is the same as a bitcon. But there is some changes, you see this graph going up here. So what is this?
29:52This is the amount of computers, amount of GPUs that are entering now. And the amount of GPUs that are participating is much higher. So what that means is that in order to break security, you need to actually break into many, many more GPUs. That's what we call this fortasia. Security has been fortified.
30:12So the economically doesn't change much, but in terms of security, it's much stronger. You need to break into many, many more computers to break into. Things become really interesting once the efficiency becomes even better when you have low overhead. So once the overhead is sufficiently low, we enter duplexia. And in duplexia, crazy things start happening. In duplexia, we see here that the amount of inference provided in the world actually increases.
30:40The token price increases, but even more so. So this is the amount of inference being provided, how much it increases. But the green here is the value of the inference. And that's even higher than the amount of inference increase. And the reason for that is because we're providing a discount on inference. So in fact, we're actually producing even more value than just the increase of inference.
31:07So to summarize, despite this common strawman approach, proof-of-use work does not make attacks cheaper. In fact, they're at least as expensive as before, and sometimes even more expensive. Second of all, on the good conditions, so if you have a two-for-one operation like Paul that has very low overhead, in fact, you're going to get a blockchain that increases the size of inference in the market.
31:35And this follows from the fact that tokens are subsidizing AI workloads. And that, on the Jevons effect, leads to even bigger market size. And that, in turn, leads to a cryptocurrency that's worth even more. And we get into this virtuous cycle where the blockchain actually is subsidizing AI, and at least a higher valuation, and a bigger substance and so on and so forth. So in essence, a good proof-of-use work blockchain leads to more social-evaluable computation in
32:08the world. That's a pretty bizarre statement, right? So we're adding this blockchain to the world, and suddenly we have more useful computation happening. Who's paying for that? Somebody's to be paying for this extra computation.
32:25And if you think about it for a minute, you realize that, actually, it's not so bizarre. This is, in fact, something called senior rush. Since we're releasing new tokens, inflationary emission tokens, those tokens are actually paying for all these new AI workloads. And this is not a new phenomenon. That happens already in Bitcoin, in fact, even the US government releases new money.
32:47But in Bitcoin, this release of new tokens, what is it used for? What is it paying for? But what is it paying for? It's paying for burning computation. Nothing. It's just paying for the burning computation.
33:05Here, it's paying for new AI workloads. And that's the amazing aspect here. So in essence, a proof-of-use work chain becomes a financing vehicle for AI workloads. As a consequence, that means that if you get a blockchain that has higher adoption or better efficiency, that leads to a bigger AI market. That means that AI provides actually incentivized to ensure higher token adoption better technologies,
33:38because that leads to bigger subsidies and, indeed, a bigger AI market and ultimately to more social useful AI computations. And that's it. Thank you.
00:00One could argue that the real currency is not money, it's energy and data. The two, what we collectively call compute, so the two scarce resources whose fusion creates intelligence is data and energy, data and compute. We believe that this dramatic shift in the production of
00:17knowledge compels us to rethink the fundamental properties purpose in the creation process of money. What money actually is? The basic scientific question that started this, the PRO project, and we, Elon and I, asked ourselves about a year and a half ago,
00:38was whether we can build a monetary currency directly from these two key scarce resources, data and compute. This is the motivation for this question, is not merely mathematical or philosophical. It is to address at least three core incompatibilities of the fiat system with emerging AI economy.
01:01The first one is that the unit economics of AI today is simply non-fungible, not all LLM tokens are born equal. Producing a deep-seek R1 LLM token on Nvidia Hopper has a very different physical cost, production cost, then producing a reasoning token, a chain of thought of Quen14B on AMD MI3s, right?
01:27So these are just two very different physical costs. The current crude market pricing of, you know, X dollars per million LLM tokens is at best a crude proxy for this actual energetic cost, but it hardly serves as a proxy for the data dimension, which is much more elusive and harder to articulate.
01:52The second drawback or limitation is that in the current economy, AI consumers are completely left out of the upside of the wealth generated by AI. So AI users like myself who are, you know, just using cloud for coding or chatbots, we don't get it in a prop or Nvidia share, well, most of us, at least,
02:16by doing so, yet consumers are the ones driving demand and model improvement through prompting RL data labeling, green for our legit and so forth. But we don't really have a stake in the wealth generated by the AI revolution, right? So that's another part.
02:37The third one and final one is that in this era where autonomous AI systems are slowly penetrating or, you know, economy and driving a lot of the value, we must have an economic system that allows machines to participate alongside humans in the same ecosystem without, you know, onboarding KYC and identification, right?
03:01That's, we just have to allow machines to somehow manage the two only resources we deposit in their hands or in their, you know, in their system, which is compute and data, right? So AI agents just receive two researchers, the data, the prompts and the compute budget.
03:21And we believe AI agents should be able to transact seamlessly and directly using these resources to manage them without any permission or regulatory or social restrictions with their, obviously not subject to, right? So this is what we often call among ourselves
03:41the money for machines, right? So it's money for machines and humans.
00:00The real protocol, the eventual protocol, which we call Pearl Gem, I think this is the probably Pearl's biggest innovation, is to follow the same idea, but in order to make the peeling of the noise the decoding cheap, maybe you can just add structure to the noise. So instead of just using brute force random Gaussian matrices, what we can do is we can just add the noise matrices which have,
00:22they look like kind of marginally random, but they actually have structures. So if we add what's called low-rank random Gaussian matrices, low-rank matrices are still end-by-end matrices if you open them up. But they have linear correlations. So actually you can decompose them into the outer product of two strip matrices. And the thing you should know is that this completely solved the peeling problem
00:46because now Alice needs to peel off those extra three products. If you do the math, you'll see that all those peeling products are going to be low-rank, rank R matrix multiplication, that's going to cost Rn squared, which is negligible compared to R cubed. Okay, so in practice, Elon and Vippo will talk about this, we'll see what's the actual overhead in practice,
01:08but that's great because now peeling is just negligible, right? But I claim this screws up another thing. This protocol is no longer secure because if Alice now chooses the all zero's matrices, Alice is left with a much easier problem than an honest player that's multiplying dense matrices. Alice now just needs to multiply low-rank matrices, rank R matrices.
01:31Alice can do this operation in Rn squared time, whereas an honest miner running actual LLM matmoles is going to pay N cubed. That's completely breaks the security. The system is no longer fair, right? There are easy instances. Does that make sense?
01:47It looks like we're trying to hold the string from both ends. We tried to make the system secure, so we added 300% overhead, and we tried to reduce the overhead, we screwed up the security, right? It seems like fundamentally this trade-off seems inherent. And I think the key innovation here, the key conceptual idea, was the following that.
02:08You know, for an honest miner, for an honest Alice that wants to run LLMs, all Alice cares about is the output, is the product of A and B, right? She doesn't care about the intermediate computations, but no one said we need to do the proof of work on the output. We can, we're free to do this on intermediate computations as well. And I think this is the key idea.
02:28So we would like to prevent Alice from exploding the low-rank structure of the noise in the intermediate space, right? And this is how we do it. So instead of just verifying or hashing the output, of A prime times B prime, which is gameable, we just saw it. This is easy.
02:48Instead, we're going to verify we're going to do the proof of work, not on the output, but on the intermediate computation that the GPU already does. So when the GPU multiplies those two R by R tiles, we will just for every such R by R tile that the GPU is going to multiply anyway, we're going to force Alice to hash the product, the result of those intermediate tiles,
03:15and check if it starts with, let's say, 19 zeros. This is exactly Bitcoin's adjustable difficulty puzzle problems. And the nice thing is that about low-rank random matrices is that marginally every R by R tile set that you see here is going to be uniformly random. This is just a R-wise independent distribution.
03:36So that's a standard property. Of course, there's tons of correlations between those tiles.
00:00Each and every proposal for useful proof of work was shown to fall short of satisfying one of those core definitions or properties of useful work. And indeed none of those attempts have succeeded. One of the most vocal votes was by Vitalik, the co-founder of Ethereum, who said that if we can find some useful computation, which is easy to verify, then crypto mining could actually become a huge boon to society, not only removing the objection that Bitcoin wastes energy, but also being societally beneficial. And then a few lines later, Vitalik adds is probably infeasible. A few years later, Ethereum switches from proof of work to proof of stake, and the rest is history.
00:00What's cool about our system, and I think you all understand it by now, you don't need any special hardware, you don't need ASICs to mine it, you can use your GPUs, the ones that you have at home, or the ones that you are renting from a data center. You can use your favorite inference engine. We have plugins for SGelang and for VLLM.
00:19So you can just do pip install and start mining basically. There's minimal setup customization right now. The existing system supports int computations. So we do need to quantize models to int right now if you want to do useful work. But we have a next generation scheme that we're working fiercely on now, and it should be ready in a couple of months supporting floating point computations.
00:46And then you won't need to do any quantization whatsoever.
00:00So I mentioned that we need to quantize models to int. In fact, we need to quantize them to int 7, which is highly weird, which is somewhat weird. And the reason why we need to quantize things to int 7 is because we're adding noise. And we need the signal plus the noise to fit into int 8.
00:16Int quantization to int 7 was not a thing. So we had to create our own quantization technique. And as usual, we came up with something that we believe is SOTA in terms of quantization quantization. We came up with a way to use existing tools, GPTQ and smooth quant, which are existing tools for quantization.
00:36But we combine them in a very non-trivial way to get an end-to-end quantization of existing models that outperforms any previous quantization, even to FP8.
TRANSCRIPTS OF PUBLIC PEARL TALKS AND CLIPS, MACHINE-TRANSCRIBED AND LIGHTLY CLEANED. THE VIDEO IS THE SOURCE OF TRUTH — LINKS ABOVE. ANY SPEAKER WHO WANTS A CORRECTION OR REMOVAL CAN REACH US VIA THE TIP LINE; WE ACT PROMPTLY.