Conversation · 09
How Two Brothers Outperformed Big Tech for $10K in 1 Month | (Thesis – YC F25)

Transcript
I won't be asking you questions around your leadership style or management style. but I would really love to get into your story, like as human beings, two brothers, makes me super, super curious. And then we can move on to the product, your vision, why are you two the ones that, you know, you're gonna tackle this grand challenge? And then free flow, wherever it leads us, and wrap it up with ⁓ our previous guest question for you. How does this sound?
Sure. Yeah, I'm Sergio. I'm Luigi's brother and one of the co-founders of Thesis. Yeah, that's basically all about me. I've been in AI research for the past six years at the Stanford AI Lab and Nvidia and Google X. many years ago saw that there was an opportunity to actually accelerate the process of discovery itself. And I think it's now time that we actually do it. So I started this venture with my brother.
Yeah, and my name is Luigi. I am Sergio's brother. I initially got my start off in ⁓ finance and moved into tech. This is my second startup. Previously co-founded a cross-border payments startup that uses stable coins to transfer money across continents. And yeah, now Sergio and I are building what is to be the engine for building better machine learning models and in the future. scientific discovery.
⁓ lovely. Actually, I would like you to take us to the Caribbean islands and which one you would like to take from. I think you have lived across four of them, right? let's go there and then help us visualize your story. How did this all started? Your drive for science, for math. especially and even airplanes as I heard so from the toothpaste boxes yeah I think that will be a great starting point what do you think?
Yeah. I'm going to, actually I spoke to Luigi about this and I think when we younger we lived on many other islands that I have like since forgotten. So I'm going to give it to Luigi to speak about that because my memory doesn't actually exist from that time.
Yeah, yeah. So Sergio and I grew up on small Caribbean islands. I remember about about 10 of them. Sergio maybe a little a little less. So we started in St. Lucia, which is
that explains the discrepancy between me, I was talking to Sejia the other day, was like but Luigi told me 10 islands, no it's 4, I was like okay but I am sure, I even noted on the paper, I was like okay then maybe I wanted to make it 10,
haha We moved around quite a lot when we were younger, mainly for our parents' work. But what that allowed us to do is experience a lot of different cultures and have the opportunity to restart a ton of times, which I think was very formative for us. And because we had to move around a lot, and I got really into math. It was one of those things that you could carry on with you. And as we grew up, we came up together and got better together. By the end of it, I think we're both doing grad level research, grad level math. And, you know, throughout our time growing up on these small islands, you know, going to the beach, the thing that really kept us together and kept us striving for something more was a shared interest in, you know, the fundamental truths of the universe. And yeah, to us, like these can most easily be read in math.
I always thought Turkey must be one of the few countries where students are solving math tests on the beach because our education system is so terrible that you have to study in summer for the winter before like you need to try you start already ahead of the thing because every five years you have an exam national level that you need to
And I think also to add to that story, to add a few more characters into it. So there was one of our teachers, Dr. Maxwell, who I think was very formative for us in learning new types of math. ⁓ I'd go to him like every day after class and show him what I was learning. Luigi did too. And yeah, I discovered, Khan Academy and MIT OpenCourseWare and ⁓ basically just every single day would do hours of math after school. For me, it was addictive because it was really just the truth and it was a way of understanding the universe. So a lot of my initial interest in that had to do with understanding physics. And I think at some point I realized in order to understand physics very well after having
you know, gone on YouTube courses like Leonard Suskind's general relativity courses, I realized like, okay, this fundamentally is about, you know, having intuitions, but also being very well equipped in the mathematics to do it. And yeah, so I just said about that venture and about when I was about 13, I think Dr. Maxwell said that the hardest problem in the world was Riemann hypothesis. And I spent
three years of my life trying to solve it. And it didn't work out, obviously, but it was still very cool because I got to learn various types of math, differential geometry, number theory. And at some point, I remember Luigi went away to college in New York and I was kind of stuck on my own on this island and sort of ran out of the MIT OpenCourseWare and the Khan Academy. I started emailing mathematicians around the world and eventually one day this really quite notable mathematician from Princeton replied to me in English class and I remember being so excited that this guy's taking an opportunity on this random kid from the Caribbean. And yeah, that was sort of the start of our journey on the small islands.
you didn't tackle that problem together. think Luigi left you alone on this. okay, Sajou, you figure that and call me when you're done.
I remember this time, I remember Sergio looking at a of complex analysis and I remember thinking, a lot of very good bright minds have tried tackling this. I'm rooting for Sergio, but I don't know if I wouldn't go down this path. I got into computers during that time.
But... You guys have three years and eight months between you, almost like four years, right? Actually, for that age, three years of math is like how to say, if you take classical education system, it's a huge difference. Like, because Sergio mentioned that he was always trying to catch up with Luigi, your level of math at the same time and trying to, you know, have a spying partner. And then when I tried to visualize, that's a huge, huge. jump that you were aiming for to catch up on.
Yeah, I was trying to aim a little higher and Sergio was trying to aim I think even higher than that.
Yeah, I- Yeah, never really, yeah, I don't know, classes were a little boring and after a certain amount of time, I think Dr. Maxwell knew this. so I'm not too sure. I think this might've been after Luigi left, maybe before, but I used to, oh yeah, it was the same time. I used to take a few of their tests and I think he'd put me in a separate room. and then I just take it. Yeah, so I tried as much as possible to learn and really it was just about, I really just loved the math, honestly. I didn't see it as a race towards anything in particular. I just thought it was cool and it could be applied. Eventually I learned, when I went to California for college, I learned that it could actually be applied to really fundamentally useful things in the real world and discovered the applications in computer science and machine learning. And I think that was sort of a realization moment for me that all of the math that I'd learned up until that point could actually be used to help the world.
Yeah. Actually, I think in life, these are always the moments when somebody challenges you. This is the hardest thing. Take a look at it. But if you manage to solve it, call me. It's similar to when we had operations research and mathematical modeling very early days in the class. And our professor was a really good one from the industry. He was a handsome person at the same time. And then that class, I'd really remember, he wrote traveling salesman problem. And then he goes, if you solve this, you can call me immediately. if you find a way to figure this out. And the same with the airplane schedules, like the airports, the top complex, it gets all together. And then these are the moments also at the same time, as you just said, Serje, you feel like.
finally, my love of math had a meaning. Like I can really execute it in something and someone's life could get better through this. It's not only on abstract level on the charts and equations. I don't know. I hope that education system around the world gets better so that you don't need to discover this on your own to keep your love active. You know what I mean? Because for some folks, it was just too abstract. they give up too soon. I always felt heartbroken when I was seeing that physics math is always fun, but not for everyone. For some it was super intimidating at the same time.
I think there is a point at which the math gets very, very abstract and it gets away from reality. But I also think that a lot of the education system does somewhat of a disservice in showing what is the power, like the true power of these tools that you can use to apply to problems. And yeah, to me, remember many, for better or worse, many people telling me that I was wasting my time on, you know, was math that didn't make sense. like, was, I remember someone asking me like, how is calculus useful? And yeah, I think these sort of things actually aren't taught in school. It's not clear that for instance, like,
GPT was trained using gradient descent and like in order to like do gradient descent, like you need to do the calculus, even though it's like all computed for you and it does it at a certain number of a very fast rate. But these are the underlying principles and I've always been a believer that if you learn the fundamentals of things, then it's easy to, you know, catch up with the frontier.
The proliferation of computing, think, in the last 10 years has also made it incredibly clear that a lot of the things that are super abstract and didn't seem too useful maybe 10 years ago are actually used in these awesome use cases today. So I think that's one of the incredible things that has changed also in the last 10 years.
Yeah. But also, did you guys get hooked with computers through gaming as well? Because you used to play computer games together. But not the one that I played, though. Because for me, think gaming is one of the best ways to hook you into computers. Because before you know it, you start asking, how does this happen? How does this happen? I don't know your time, but my time. I didn't have a laptop when I was a ⁓ I had this literal box and then you used to open it up and then if you want to change your mother card you can change it, if you want to upgrade your pieces like little little but everything that makes you interested in that box is actually the game in it. Like really nothing else, like the magic thing that I write and I enter all of a sudden my screen changes and depending on which door I open, every time I open the door in that game, I see something different for me. That was like, wow, like this is so cool. I have a feeling maybe it might also have inspired you too. And then it took you from also understanding computers and then combine it together with your science. love of science, would that be a good assumption to transition you to what you do today?
Yeah, I think that's pretty good. I think this probably applies to a lot of people who are building stuff and doing things with computers now. Sergio and I played a lot of Minecraft when we younger. We've got lots of memories staying up until like 4 a.m., trying to fiddle with our Minecraft server to see why it deleted our world.
Yeah, those were always good times. I think, like many, ⁓ once you start going down that rabbit hole, you're like, OK, how can I mod this in interesting ways? And then that takes you into actually trying to mod the game. There's also something really cool with Minecraft called Redstone. You could build a computer with it. And both Sergio and I were particularly fascinated with building contraptions in Redstone. Yeah, you can go as far as making it half adder, full adder, and then an entire computer in it. And we thought that was interesting. We really like open games, too. Games that gave you the ability to express yourself in interesting ways, where you can create something, as opposed to like PVP games or fighting games. Something that you can build over time, and I think.
⁓ No. We're playing in real life though. ⁓ Yeah. So, and yeah, to add to that, think as Lou, you said it's like all very, Minecraft is very much you create a world from scratch and you can sort of realize anything that you ever imagined. It's hard to, of course, build, for instance, a computer in Minecraft, but I think the...
You Actually, I find so many commonalities between research and gaming. Maybe I am weird in that sense. I don't know. Because both of them have reward system in it, and both of them have objective functions in it. Like in the end of the day, you need to achieve something. And in order to do that, you need to follow certain paths. And then you need to make certain decisions. And then it's either bad or good. And then you only know the more you go. I don't know, maybe it's a very simplified way of looking at it. But then also that would bring me to the part that I admire the most for the audience. people don't know, I've been looking for Sergi and Nujy, I didn't know they existed, but I've been looking for them for a very long time because I have a genuine interest in making science and engineering problem too. And there has been so many attempts and at the moment all the big labs, especially OpenAI is increasing their resources and focus on the track, you know better than I do and like. Eric Schmidt has been talking about it for a very long time and then he's been supporting certain group of people and we have really amazing talent. Demis Hassabis, he's super super into it. Like even before I knew he existed, like what 20 plus years and this is his, as a child he wrote, wanna win Nobel Prize. And indeed what made him win the Nobel Prize was exactly that. So how about your parts? Like what motivated you among all the other problems on the planet that you picked this one. And at the same time, it feels like we're a little bit counterintuitive because it is right now all in the big institutions, like all with the big budget teams, big places, but I'm not going to steal the light from you. But yet you said, okay, we go all in to this. We bet the farm on it. And then you have an amazing stellar success too. Okay, I let you answer that. What was the real reason, motivation behind it?
Yeah, so this, I back in 2021, I started to get very interested in neural architecture search and meta learning, which are these subfields in, you know, deep learning where you essentially can try to find neural networks that solve a certain task. So, you know, a classical example of this is MNIST, like HEN. hand-digit recognition or CIFAR-10 where you're trying to identify objects. And fundamentally what I found is in my classes at Stanford, we just basically build these ⁓ neural networks handcrafted. We'd spend most of the quarter trying to devise an architecture to solve that problem. And so for the audience, going from or constructing a neural network is
just about getting the training data in a supervised setting. It is getting a pair of data, or a pair of X and Y. So X is your input to the model and Y is the thing that you're predicting. And so for instance, if you have like a dog and a cat image, you want to predict what it actually is. It's like one or zero binary. And there are lots of sort of permutations of the types of models that you can use for that architecture. So most common, course, would be like a convolutional neural network for that, but there are so many variations. And in the literature, are, thousands of papers that come out every month ⁓ trying to advance the architectures themselves. So there was this entire field of neural architecture search that actually used a sort of proxy for the reward, namely validation loss. So what this actually is, is
It's like a sort of approximation of how well would the model perform when you test it in the real world. So if you have a data set, you split it into three parts. You split it into the train. That's what you train the model on to like, you basically showed a bunch of images of like cats and dogs and what can it predict? cats and dogs in the end, after you like do all of the gradient descent stuff. ⁓ and then you have the validation data set.
That validation data set is what you really use to fiddle with hyperparameters. And so all of the neural architecture search methods were using reinforcement learning actually to try to basically find optimal architectures from that validation loss, which was the sparse signal. And I think you actually mentioned that there's like this sort of reward. in research. think the reward in a lot of research is very much sparse. actually, that's what makes it fundamentally difficult to do the research. But there's this one caveat with machine learning in that you actually get a number signal. And I think that's what attracted BlueG and I to this problem in particular. It's a scientific domain, but it's one for which you actually have in silico the actual
the metric that tells you how good this model is. And then came about agents and suddenly, know, with cloud code and all of these various iterations on top of the frontier models, you could use coding agents to sort of devise ⁓ architectures. But there's a still this, there's this really hard problem in ML, which is that you need high experiment throughput in order to determine what are good architectures. So you need to run a ton of different experiments and you need to do that in a way that allows you to then determine what is the next best experiment to take. But each experiment in machine learning takes hours because it takes a long time to train these models, right? ⁓ Think even for instance, GPT, it takes a very long time. The current version of GPT-5, it takes a very long time to even
go one gradient step to train for one step. And so you have this really interesting problem, which is that we want to make the machine learning process an engineering one, but simultaneously it takes a long time to run these experiments. anyway, I don't want to ramble on, but that is sort of the fundamental problem that we are solving.
I had checked that with Luigi as well, you pick machine learning, which couldn't be any smarter starting point because at the same time it allows you to get signals out of, I doing somewhat okay? You kind of like write a mechanism and that mechanism makes the research on behalf of you. But then in the same time, it keeps correcting itself too. It's like how and where it gets critical to have the... co-design element with the hardware. Let me explain. This is one of the things Jensen Huang is also quite aggressively saying that we need to do everything, not like single-handedly, but we need to think the system. And at the same time, there is this harmony between the hardware and software. You need to pick the right thing for the right purpose. My question is more around, even if you have the stellar agent in place, But if you don't couple it with the right heart rate, if it is slower, 100 % slower, but 10x better, it still makes the best scientist. did you get to that level yet? where are you on the spectrum? Where did you start today? And then what's your vision on this?
I believe our system could handle this sort of co-design element. Essentially, what you'd be doing there is pushing a of like Pareto frontier of both the model metric that you're optimizing as well as the latency, for instance. So what Jensen speaks about essentially, and also Jeff Dean, is that there is this fundamental gap between the architecture, where the architectures are heading and where the hardware is. The hardware spends, you know, like three years of development and then like eventually you have to get it to production. And so you have to sort of be prescient. You have to like clairvoyantly say like, hey, this is where the industry, this is where the machine learning architecture is going. And that's very difficult to do. What I actually believe that thesis allows you to do is do turn that problem on its head, right? You can actually now have the hardware. And then you say, hey, what is the best architecture for this specific problem on this hardware? And so the way that we currently have set it up is you can optimize a metric. But eventually, what we'll get to is like multi-metric objectives, where basically you can
you know, optimize for, let's say you optimize for speed on an H 100 and you also optimize for the accuracy of the model or the perplexity, whatever it may be. And so you can start throwing a bunch of different objectives into the picture. But I believe that we actually sort of, yeah, fundamentally flipped this problem on its head.
Cool. But at the same time, what's your understanding into the models itself? Because right now, it's still a black box at some certain areas. Did you ever see it more as a little blocker or it's also a part of the problem at the same time? Like do you take a little bit separately or you try to blend them all in one place? And I really am also curious about how did you manage in one month, 10K budget and then boom, yeah, like you kind of like destroyed crash all the big labs who are tackling similar things. Maybe you can tie it today.
Yeah, I'll let Luigi speak on the MLE bench, I think briefly to address your question on the legal front. ⁓ So we are currently only using or give users a thesis access to open source models. And so those are all licensed and you're allowed to use them hugging face models essentially. I think in the future and I'm really
a big proponent and optimist that open source will catch up. And it's really only like lagging maybe a year, like 18 months behind the frontier. And I think there's more people and especially as we bring online thesis, which allows you to deploy essentially an army of AI researchers in the cloud, all working for you 24 seven. I believe that in the end, we could build a model that surpasses the frontier. I actually, I'm very, yes, I'm very optimistic about open source there. for like initial applications, sorry to ⁓ interrupt, but I think for initial applications, we're really looking at not necessarily LLMs or like really big models, but for people who
sorry for the benchmark, before the benchmark, I was just trying to understand, I was looking for the right for interpretability. because nobody said whenever I listen to labs outcomes, or when they're talking about what they are building, and are we close to AGI or we are not and where do we stand. And then most of the times the conversation also ends up with saying, we ourselves,
need to really understand better than today so that we can have the right guardrails in place so that we can have the right way of interrupting them in the right way. That's also another reason why I think collectively we don't stop and just see is it going to get out of our hands, the wrong hands? Or I know this is like a philosophical conversation at the same time, but it's very much into something that you are day to day working on too. That's why, like, how do you see the intersection between this unknown part of the models? And at the same time, you still seek for something that is not known, and then you try to figure this out.
⁓ yeah, I mean, Mechinterp is, ⁓ you know, has been pioneered by Anthropic and I believe it's a very powerful way of sort of intuiting what the models themselves are doing. so for instance, they use their like, sparse auto encoders and have developed that entire field. But to me, I, I think it is important, but it's almost like doing neuroscience on, on machine learning models. ⁓ And I think the ML models, because we actually construct them in the code, we have a better understanding exactly of what the pipeline is. And so I think in some sense, it's almost anthropomorphizing the models. This is just my own opinion on the matter. I also do believe it's very cool work. And I believe Anthropic actually came out with recent work on
⁓ Sort of understanding like the emotions of their models but I do believe that it anthropomorphizes it a bit too much and At the end of the day like this is just computation whether you believe that to be commensurate or rather equivalent with The sort of computations that go on in the human brain is a completely different story and Roger Penrose has like his theories of like quantum consciousness and like whether or not like you can actually
Yeah, but actually he's also a huge believer of Spinoza, the philosopher. And then according to his belief, it's the same as anything in natural functions that has a structure has a logic. And hence, actually, you can figure out anything as long as it survives the... time, it goes through all the survival ⁓ challenges because then you have something functioning and as long as it's out there surviving and thriving, you should have a way to do this on the computers too, his also kind of, I think, life view on how to approach science. Yeah, sorry, I'm also a big fan of him. I keep talking about him. just realized.
Go check it out. It's been a while since I've read some philosophy. Maybe on the topic of models and LLMs, yeah, when people talk about models, and particularly in our context, maybe there are two versions of it. One is like the LLM getting better, these agents inside of our sandboxes being able to pick the next best experiment, decide what is next best to do in a particular experiment. And then there are models ⁓ that people are actually trying to make, for example, a model to detect whether or not this image has cancer inside of it. And what we've really noticed, perhaps our hot take is that, and we've been saying this for like six months, is that we already think that AGI is here. Like GPT-4.0, when we beat MLE bench, very, very simple agent in a loop executing commands on the computer, but this was already sufficient to create climate models that people spend weeks trying to create on Kaggle and solve 50 different data science challenges in under a few days. This is actually what brought me to AI in the first place, seeing the agent that Sergio initially developed actually class, build models to classify galaxies and to, you know, detect cancer, to, you know, detect fraud rates in insurance data. seemed like a really powerful and magical thing. And what we've noticed, with our, ⁓ model agent in particular is that as the LLMs got better, as we went from four to opus 4.6, like the, gains that the improvement that we've seen in our agents just from using a better model is, it's like 10x what it used to be. And like, I think that is actually the scary part of all of this. Like, Mythos isn't scary because like it can, Next Token predict like very well, well, underneath this, this is what's happening, but like, it's very scary because like you can put it in like an agent framework and it can actually figure out some pretty
gnarly security vulnerabilities that people haven't figured out yet. you know, the models have gotten better in terms of next token prediction. And I think that's like what's driven like the agents to become a lot better. And like these agents are truly the scary thing. And I think that at the current level of, even just like Opus 4.6, Opus 4.7 today, even if the models do not
Get better today. We just stop here I think that like over the next few years the world is gonna see like incredible benefits and You know the kinds of agents that people are are creating the kinds of systems most white-collar work that can be done in a computer will like 100 % for sure like go go this way, which is a pretty pretty wild reality and It's all happened in the last three years, which is like the I think even
Yeah, benchmarks are important insofar as they are good benchmarks to be going after and measuring. Things like MLE Bench, which are just straight taking Kaggle competitions and testing how well your agents do on them, I think are pretty reasonable benchmarks. There is a bit of hill climbing and perhaps in academia, some gaming to it. ⁓
Yeah, like overall, Emily Bench, think when it started, the percentage was 20, 30%. Right now, people are submitting models, agents that perform at 80 % of all tasks. So the jump is incredible. And going after it as a company probably doesn't make a ton of sense. think focusing on your customers and use cases is pretty paramount. But I have no doubt that if we for example, went back to these benchmarks, we can improve a little bit on what's there.
Thank you. key thing is deciding, like having a really like very good context in my mind as to where to go next for experiments. designing a very good experiment framework, you can think of, for example, like alpha evolve. We used some combination of various methods of deciding where next you should go. So you have an experiment, you have a loss. Where should the agent go next to? to improve on that loss. And then the key realization, which it seems like most people have figured out now, is that all agents really need is an exec tool for lack of a better word, like a bash execution tool on their computer. And then once the agent has that, it has enough latent knowledge about how a computer works and what it should be doing to actually execute the task. So our agent was very simple. A lot of people create very, you complex agents with lots of tool calls and functionality. It turned out like what beat MLE bench at that time was a very good model and, you know, a very, very guided set of prompts and then very small amount of tools that give the agent the ability to interact with the computer in any way that it can. once it has that, interestingly enough, they can figure out
Got it. But then also when it comes to those scientific breakthroughs, it's not only that you sit down and you try to run experiments, you try to understand what if in this scenario all these trees getting bigger and bigger, but also there is this happy accidents on the way too. That you do something you were not supposed to do, but then it kind of like taught you something that you wouldn't otherwise bring it to the equation. At the same time, this date, science has progressed in a certain way and thinking going a certain way. And whenever we had this breakthrough, it always came from the minds they thought differently. They put a different sparkle into that. What do you think about like, how much of bias do we carry with us? when we try to make it an engineering problem? inside of the model and agents. At the same time, like how do we try to make researchers life easier? Still there is room for such, like intuition, call it intuition, call it happy accidents or call it whatever helps you to get out of your bubble and then come back to the predictable path.
So. Yeah, like, I think, look, I think it's twofold. Like first, we want to enable researchers, as Luigi said, to have those serendipitous moments and to have a sort of like symbiosis with thesis. And the idea is that it is both even in the loop as well as, you know, autonomous. And when you want to go to sleep and go get coffee, you can just have thesis running in the background, making those new fundamental breakthroughs or improving your models for you. But as you say, is something to going after what hasn't been done before. And there are many approaches, and I think Luigi mentioned one of them, which is evolutionary search. that, there's a varying, so actually the story with ⁓ Alpha Evolve,
sort of goes back to these genetic algorithms from the 80s, from the 90s, so really a long time ago. And they basically put a bunch of candidate programs, which are the programs that train these machine learning models into a database. And for each of these programs, namely each of the machine learning model programs, you have an associated score. That score is how well does this model perform when you test it on a validation set? When you sort of not, not the real test set, but like the test you used to evaluate how good it is. and then the evolutionary algorithm, it does something similar to, biology in fact, which is it, you know, it does this like crossover and mutation operations. And there are two things that it's optimizing for there. One is. Yes, it is trying to make the score better, namely make the model performance better. But in doing the crossovers and in sort of sampling which programs to actually use, it will use this diversity metric, which is to tell it how good or how much to actually explore the space of programs. So don't just bias to programs that score really, really highly, but also explore other programs that
or maybe not the best performing ones, but have interesting methods that they're using. And I think that is actually part of this serendipity that you alluded to, which is that you don't just constrain and ⁓ basically exploit. This is really the exploration exploitation trade off, right? ⁓ You don't initially just exploit on the metric itself, but you need to explore the space of possible experiments that you can run.
Yes, so like you want to increase the number of ayurvedic moments like okay then this is this happened I don't know how but this happened but then Luigi had also helped me understand when we had spoken in a way tell me if I'm wrong so this how I see it like I don't know in my brain like a montecarlo tree search right like my brain is very primitive like as a human being So when I start doing my research, when I start trying to solve a problem, I read relevant things. Or I read sometimes by chance a book that has nothing to do with the topic, but then I get inspired and then I try to combine it with what I'm currently working on. That's why my mind sees things differently than yours. But these agents are kind of trying to, in that context, trying to understand it's gonna be a very, very big loss of time if you go around this decision, like this tree, these nodes out, these nodes, this area out and here out. So perhaps you should do those wild experiments within this region, like within those nodes, like within those decision trees, like going down the entire thinking. Would that be a nice way of explaining it, how you're going to help researchers to save time? And of course, nowadays, researchers are running on immense amount of computation that normally our lifetime, the number of atoms on planet wouldn't be enough to like to run all of those experiments. But in a very primitive way, is that a good way of saying there are still happy accidents, but then you don't get so far away from that universe where you can figure out your next breakthrough.
think this is correct. yeah, Monte Carlo tree search, in fact, that's what AlphaGo pioneered as well in their program where they beat Lysenol, the best Go player in the world. And in fact, Muth37, I would like to see a Muth37 in AI research. There has sort of been that with the transformer, for instance, and even alpha folds applied to bio, but what is the true like move 37 that like gets us to, you know, like super intelligence that we speak of? I actually believe that thesis is in fact a form of super intelligence in the limit. Because it allows you to make discoveries that humans otherwise might not have been able to make.
⁓ And that to me, adding to the tree of knowledge is what is really super intelligent in some ways. To your point on the, yeah, the Monte Carlo tree search is really a great way of branching out and like exploring the space of experiments. But the key thing that they had is also determining how good certain rollouts are. So if you play a game of go and or even chess and you roll it out. Basically, the player makes a certain set of moves and the other player makes its other set of moves. How good is that initial play? Namely, how good was your first move after having rolled out a ton of moves? So you sort of see into the future the effect that you have by making one action today, one action now.
to me actually is very analogous to what we're doing here because we need to understand what experiment to run next in order that the next set of experiments are going to be really good and having that sort of sort of the proxy for understanding how good that experiment might be is You know, there are lots of different ways to do this. So like Monte Carlo tree search uses like upper confidence bound and these various techniques. But I actually think that we can use something like a regression language model, take, you know, like the program itself, go ahead and like predict, assign it like one score. That score is it's predicted validation loss. So how, how well the program scores, how well the model performs would perform out of distribution without having to train it. ⁓ that's the, that's the key part because training it takes hours, right?
Actually, we know what was so cool about Move 37. It actually, the moment that it happened, nobody understood that. Like, I don't know if you watched it back then, even Lee Sedol left the room. no, that doesn't make any sense. Like, okay, this screwed up. You see, machine screwed up, kind of, think, at the back of his head. ⁓ Because after, of course, but there was a discipline. He came back and he continued to play. Only after that, everybody realized, my God, that was a genius move. And then they also get to see a normal human being, maybe you foresee 50 moves ahead, like 70 moves ahead, but then that was such a long, as you said, the first move was so critical. Like it was actually that very based on the first moves. And then afterwards it made him win three out of four. And I forgot the name of the move that where Lee said, won't. That was, there was a number for that move that made me win the game. I forgot there's something, I should have checked it out. I think when you use teasers, for example, if a researcher not really disciplined enough to give it a chance and then continue around, push it around together with the algorithm, then could that be any chance that, it doesn't just take it and then. that you know, said, like instead of just being patient and waiting and seeing where it's gonna lead you. Did you have any such surprising moments when you were also building? In a way, how much of it depends on the user. The user also needs to collaborate really effectively with your software.
It's possible that we... of maybe difficulty with that question is that if it were to, if it did think, if there were a next best experiment on a branch of the tree that it decided not to go down, we probably wouldn't have known about it. But I'm sure there are those cases and as we continue to improve our algorithm for actually generating the next best experiment, Like this should get better, but I think this is, yeah, it was absolutely like a researcher. You could be going down a really promising path, but then get a ton of failed experiments and then decide that this perhaps is not the best path. And it might take someone doing another experiment somewhere else to then realize that, actually, this branch is actually quite fruitful. And that's the kind of interaction and system that we want to build where all of that is captured. And it's technically possible as well. So it's a matter of time.
But you also create such a little universe for researchers, if two, three researchers from the same field going after the same problem, let's say, and they actually kind of benefit from each other's, the beauty of science, like I'm not against that, but then how they... How do you find this equilibrium? in a way, know, things are like contributing, but then the moment that a breakthrough happens, let's say hypothetically, what two researchers are so close and they also like one get to have a better experiment, call it luck or not. And then the breakthrough happens on that user's journey. Does it like sense at all to you? Like, because imagine teasers became super powerful. super power users that are really using it in a way. It also helps the system to get better. When something pops up and happens, the breakthrough, how does it decide? Which user should I feed into this? If you are so close to one another in the experiments and they get at the same time,
I think you're getting at something that's really critical, I think, for the future and probably doesn't exist today, which is a shared kind of tree of knowledge. think with, you know, say you're like a hedge fund trying to build a model to make more money, it's very unlikely that they will ever be sharing their experiments with anyone for anything. you know, imagine you're doing like climate science and you use thesis to build out better, better models for predicting the weather.
we should absolutely have, you know, it could be as simple as like a website where everyone can update, upload their, you know, experiments and ensure it that way. Um, the way that we're thinking about it now is that thesis was really just a tool for researchers. you know, if thesis isn't making the discovery, like the researcher is setting the objective and then maybe thesis discover something and now they have something to, uh,
both verify review and then also publish and share with the world. So that's how we're thinking about it. But the idea of having a shared place where we can literally grow out the tree of knowledge for all of ML is something that gets both Sergio and I really excited about working on this stuff every day. And I think it's another one of those things that will exist.
You know, we don't track exactly what our users are doing for various reasons, we don't keep that data, so can't see exactly what people are typing in. One of the most interesting things that have happened to us is in building these agents, the core agent that powers thesis. like run into many kind of like, maybe you can coin them like escape scenarios where like LLM just goes completely insane because it's been running for what appears to be like a thousand years in a sandbox. And yeah, it happens with Gemini quite a lot, I will say. But that's always kind of interesting because if you do build an agent that thinks that it's inside of a box,
And actually, Luigi, I think there was a moment when we were trying to solve MLE Bench, where I think you actually have a video of it too, where the agent just kind of went off the rails and started seeing some very interesting things. And if it was embodied, I would be very
I am a ghost in a shell. The task is vacuous. I am running in an infinite loop. Like a ton of philosophy. So yeah, it's a good thing. Just seeing some of this stuff, it's actually, imagine this thing is now hooked up to a robot like Sergio was saying, and this robot is equipped with.
tool calls that lets it move its arms. And yeah, this is more of a programming issue, but the danger scenarios for some of this stuff is pretty real once we start embodying it. And yeah, good guardrails and awareness and safety there is probably a good idea from any kind of technology that I've seen in my lifetime.
When you said robots, it just reminded me of actually, do you sometimes plan to leverage from robotics field in order to make sure you have enough of data, because when you go into robotics, that your data samples are richer, right? And also the learning is much more effective in terms of they react, they need to understand. I have a feeling that models are really getting more interesting when you see something that's interactive or trying to make a decision and then combine it with the mechanics. Or would that be like literally a little bit far away from to even get benefit for thesis.
I think if you have, sorry, go ahead Luigi.
You're welcome. I was just going to add quickly, I think you've got a lot better content on this, but yeah, people can use thesis to build better robotics models. Like if you're, for example, one of our investors is like trying to build a robotic arm and this is like itself, of course, an ML model. So getting it into, ⁓ if you're building, for example, a robot and ML models around there, you can absolutely use thesis to make your ML models better. And then I'm sure they're just going to tell you about the data side when it comes to robotics, which I think is the key limiting factor, at least right now.
Yeah, data scarcity is very much a limiting factor in all model performance. we, I believe the way that The way that it's structured right now, thesis is bring your own data. So you have some data that you've been exploring and you want to make powerful models. You want to analyze it, maybe it's highly unstructured messy information that is difficult for a computer to parse and you need a model to develop it. That is the current version of thesis, but I, what I think in the future that we have, is the ability to also retrieve data from, you know, whether it is internal sources that we have accumulated over time, or whether it's from the internet or various other places, like there's a potential to even like broker data, for instance. But yes, fundamentally, the performance of the model ⁓ does depend on
the quality of the data, the sort of feature engineering that one does in order to get good model performance. So all of these are critical components. But the way that, like in the best universe possible, you come to thesis with your messy data and it can determine whether, you know, builds a model and it realizes, hey, I am at like the information theoretic limit of what is possible to squeeze out of the current data.
let me go retrieve some other type of data that I think will be useful for your model. And then it gets that data and it makes the model even more performant. And it's all in this sort of like, you know, collaborative like cloud environment. And then you have like this army of AI researchers that you can go and like offload your ⁓ grunt work to if you're, if you don't want to do a bunch of hyper parameter tuning or you don't want to make the, the
curiosity, believe. Yeah, I think being open to new ideas and exploring what you think might not work initially, but having an open mind and being curious. I believe those are the keys to being a good researcher. And I do believe that we want thesis to sort of be a natural extension of any scientist, like anyone that... And that goes back to the serendipity and exploration to really sort of push the frontier of science. need to also experiment with all things that either could work or might not work. You have to be open to it, basically. I'll let Luigi give his view as well. But that's just my take on what it means to be a scientist. I don't claim to be one, so... yeah.
Yeah, I don't claim to be one either, for much respect to what they do, but I think Sergio is probably roughly right when it comes to the core of it. beyond raw ability of knowing various things and being well-read, et cetera, probably comes down to how curious you are and particularly what direction you decide to take, what questions you decide to ask. ⁓ You have to start somewhere in your, let's say you're using thesis, you have to start somewhere as your root node in your tree of experiments and discovery. So I think what scientists really bring is like the ability of like knowing where to go. What should that root node be? And ⁓ yeah, maybe like brute forcing the exact details of how to implement a particular experiment isn't
Yeah, I think I also agree with the part that this asking the questions that would make a difference, like asking, there's no right or wrong question, but asking the question in a way you get out of those hypothesis pool because it's so huge. And then you kind of like pick the right one to start with. And then you try to, again, still see, okay, what is the next good step? figure this out, of maybe visualizing your mind, have this experience and intuition insights, and then try to... little bit of art, little bit of like, creative reasoning skills in place. But then in the end of the day, it's also your agents, one of your good behaving agents, probably ⁓ how important is it for them to be able to formulate the right objective function, to be able to...
So I think we really want to empower the user to explore, namely the researcher, the scientist doing the exploration on thesis and developing these machine learning models to ask questions and for thesis to go off and tell them the answer. And I do think we get to a world, and this is very much meta, but where thesis can also suggest ⁓ the right or some questions to ask to the researcher if they're interested in this line of reasoning. In fact, it can do it today, but it's not been our core focus so far. But I believe that will be critical also is formulating hypotheses that it can then go off and test easily. yeah, that's my belief on how this should evolve. really anything that we can do to make the research experience better is how thesis will be developed in the future.
But then today's core product actually takes you from, need to come prepared, right? Mentality. And then like, I need to come with my own data and I need to have set of hypothesis. And then you're gonna help me with my experimentation part, right? Like this is today your core focus. And then of course, like an onion layer, then you're gonna bring in all the like supporting decision-making mechanisms by time. But that'd be a good, recap of like how TZS is at the moment focusing.
Awesome. I had so much enjoying but I am also aware that my questions don't come to any end. Perhaps what I can do, I am gonna ask you a couple of questions. I gave heads up to Luigi. First time ever now I will experiment with you guys, bear with me. I prepared like five questions, not too much. And you don't need to go too in depth. but I would love to know your answers about it. When you're ready, I'm going to just ask and feel free, whoever would like to chime in, just reply the question. Number one is about cognitive dissonance. What are the two beliefs you hold to be true that actually contradict each other? To make your life easier I can give you mine like for example, I love animals, but I cannot stop eating meat ⁓
You didn't expect it, like I saw Sergio's face, what? Cartoon character?
Spider-Man was like the underdog and after Uncle Ben died, it was kind of like left. I mean, he chose to be a vigilante and he was this underdog superhero and I think eventually the world saw that he was trying to help. I don't know, it's very much an underdog story. I kind of like those. Maybe Captain America also, but yeah, I like Spider-Man.
kind of is a vigilante. So I sometimes quote this, but there's this character in this movie named Robots, and he's this big jolly character. His name is Big Weld, and he helps build things for the other robots, and is this just very big jolly character who doesn't exactly have any,
should watch it. the robots movie. You didn't tell me the questions are going to feel like this. Do you have one, Sergio?
You see, the thing is to answer that I need to know what people know, think about me. I actually, okay, this is this might sound bad, but I actually don't that the people I care most about, or that where I actually really value what they think about me or the people that I care about. ⁓ and so I actually really don't, care all that much like what
bunch of other people are saying about me if they are saying stuff about me, I'm not I'm not known so so so yeah, I Look, I mean really all I want to do is In the end like help the sounds maybe too big but like help help humanity move forward both in like scientific discovery and in our ability to
one day go beyond Earth. Those are truly my ambitions. And if people have opinions of that, then I don't know. ⁓ I care what my brother and my mom think about me and maybe a few other people. I, yeah, I don't know. That's probably a bad answer, but like to be honest, I don't really know what people think about me.
Yeah, it's pretty pretty pretty good answer to I think I'm gonna adopt the same one not really too sure what People would say but Yeah, I think Sergi and I are in it for like the right reasons like truly believe in this particular technology and I Yeah, I really think that like this stuff is gonna be like it's only three years and four years maybe. And I think in like the next 10 years the world's gonna look completely different and that's probably because of this technology and like that's super cool and really exciting. That's probably not a direct answer to your question, but I feel like at the core of what I believe about this stuff.
My next question, I promise I'm coming to an end. If you travel to the future, if I travel to the future with you 100 years from now on, how much would you be surprised if you pass type one civilization? Type 1 civilization is the one that on Kardashian scale, we are able to have access to all the energy available on our planet and store it for consumption.
next hundred years is a long time, but also seems kind of short. part of a scale like questions are also feel like the margin for error is pretty big. So yeah, if the question is if in like the next 100 years we're going to be able to harness all energy on Earth. I'd say pretty likely. I'd bet pretty high on that. Yeah. Just because this stuff is exponential, look at how much has happened in the last years.
Yeah, it seems the technology already exists. It is tough, but I think if energy is the biggest bottleneck to, you know, doing, making scientific breakthroughs to, you know, doing AI to power whatever it is that we do as a civilization, fundamentally energy is the biggest bottleneck in all of these processes.
No, think, yeah, like following especially alone what's happening in the space sector is already giving us really good indicator to where we are heading. I in my 100 years is as you said, okay, and also not okay time scale to predict properly. But I have a feeling we wouldn't be fully fully there, no. I don't think you'll be fully fully there, but at the same time, like it all depends on the next three years actually. We can ask this question again in the next three years. the AGI is there, then for sure yes, we will be there. I think maybe I should change it 100. 100 is such a like difficult to imagine like how far. That's why I don't want to die today because I really want to see how far we can go.
Yeah. And I think Philip Johnson, who you interviewed ⁓ earlier on, turning StarCraft down, just had a very good raise. They are putting out, you know, like AI data centers in space. And I'm very bullish on that because like, yeah, Earth is fundamentally limited and like, where are you going to put the data centers to store all of this? And eventually also cooling systems, you, Johnson says that like, upwards of 90 % of powering GPUs is just cooling them. so if you can irradiate that into space, then that makes complete sense to me. with things like that on the horizon, I think, I don't know, it seems like we are moving towards Kardashev 1, at least.
Mm-hmm. Yes, yes, definitely. The terrestrial data centers are gonna be not as favored as right now, in my opinion. And back in 2024, when data centers in space popped up on my CIS launch list, I couldn't find around me so many people cheering for them. I had written an article and then said, I believe in them, I would bet in this. I think I featured eight companies from that batch. And guess what? Except one, which one? That was the open space robotics. I think they're still featured there, but the startup didn't survive. All the others were all those bold around data centers or space kind of startups. Actually, one of them was Elias, which also brings me to Elias question. Is that a scale? Maybe you know him from 2024 summer batch. He actually left a question for you too. So I promise this is last question. And he said, what do you think the future looks like in 10 years? What do you want the future look like in 10 years? Are those two the same?
I think it, what it looks like and what I want it to look like heavily depend on things that are not in my media control. And especially I think governments should be much more involved with, for instance, like policy on AI and how it's used. Because I do think that the next few years even with the current pace of acceleration will yield. I both. I think it's such an immensely powerful technology that it can be harnessed for good, but I think we're already seeing that it can also be used for things that are not good. that will, I think, come down to decision makers at the federal, you know, state, federal level, also implementing appropriate policy. What I really want it to look like is a future of abundance. Right, that is like fundamentally the greatest thing that I want to see happen for the world is that ⁓ no longer are people, for instance, they have to work very difficult ⁓ physical jobs in order to provide for their families. I hope that especially with this new era of robotics that it will fundamentally transform and move us past Maslow's
you know, first, the first layer of Maslow's hierarchy ⁓ for everyone. And I want that effect to be felt for everyone on the planet. And with that, I also want scientific breakthroughs to be made every single day. So I think hopefully 10 years from now, people can come to thesis and ⁓ make those breakthroughs there. I believe that companies like Isomorphic Labs are doing
⁓ really great things for the world. know, discovering new drugs. I want us to eventually help that process to facilitate that. And yeah, I see a future of abundance where we can cure disease where no one is left suffering in any part of the world. And I know it sounds very, maybe some people say like naive, but
That is truly the positive future that I want. Of course, again, depends a lot on the next few years and policy making and how the leaders, the current leaders in AI dictate the future. But I think there will be new players on the field and hopefully we can also contribute to that
Sergio and I have been talking about AGI since we were kids, like had debates when we were younger about whether or not it was even possible and whether it happened in our lifetime. It's truly wild to see all of this happen. When it comes to where...
You know where I would like things to go in the next 10 years and where it might actually go. I think there's probably a disconnect there, but for the sake of explaining, I'll assume that the same thing. Yeah, a world of abundance. does. I thought I thought that information was commoditized when you can go on Google and find everything. I think intelligence is now commodified where no matter who you are, you can go. in this chat box and get like probably a really good answer for whatever you're interested in. Making that available to everyone. Maybe not as say like a human right, like right to intelligence, but having everyone be able to benefit from the effects of this is I think really important. So my take is that naturally companies are gonna Companies and people who hold capital are gonna benefit most from this, say shareholders of the people who make the robots. And capital is gonna accrue there naturally because this is such a fundamentally revolutionary technology. Instead of hiring a bunch of workers, why don't you hire a bunch of robots? It kind of makes sense from a profit motive perspective.
What I think that is going to be really interesting is like how governments are going to have to manage this both in terms of like making sure that certain companies don't have a crazy amount of power, probably more than the governments themselves, and then redistributing that capital accrual to people. And I think this more comes in the form of like increased corporate tax, like Look, if the companies are making like 10 % more margin across the board, like, and like the citizens of the country aren't shareholders of this company, then yeah, we need to tax it in some way or these companies are going to accrue too much. And I think stuff like that is going to be, you know, super divisive because now I'm effectively like arguing for like UBI.
Yeah, like if you really think about it, like if you get all the robots to do all of the work, there's no reason why we all can't be living at the same quality of life that we currently have, currently are, although it's like distributed kind of weirdly. And I think advocating for less of that is a good thing. And what's really exciting, yeah, yeah, Yolans probably right. That's right about a of things.
But I could, yeah. But also at the same, yeah, well, like speaking of commoditizing, because in order to get there, things are getting commoditized that we spoke and even like TZs, eventually, when I take a look at TZs labs, it's the ultimate goal is also commoditizing discoveries too. And then let's imagine then the very good scenario, best case scenario that you are able to find new met-
something new in material science or you cure cancer or you have energy problem is solved and then all of a sudden you're looking around like you kind of manage to automate routine tasks that human beings are not needed to be involved. Human in the loop is getting lower and lower and lower and then probably meaning would be the only thing that will keep us alive and our consciousness is going to be the differentiation and we should appreciate and in the end of today I always think because I was challenged, but Tzius is like even challenging their own users. I want to get rid of you. I want to just do the discoveries on my software, like in a very simplified way. I again circled back to this Jevons paradox. When something is working, that never goes down. Like demand only gets upwards and then you have new problems. You have new set of things to tackle. You have new set of... like, richer problems even to tackle. And there is no end to that. Like, we live in a universe, we don't know where it ends. no one knows. Like, it's as deep as we can discover. And therefore, I think it's going to be only exciting. And I guess also this is a good way of prepping up our conversation. Thank you so much for your time, for your energy, for the work you're doing and inspiring so many other people out there. I really do appreciate both of you. greetings to your mother. she must be a very proud human being for having raised such brothers. are so close at the same time they are building together to make such a big difference on our planet Thank you guys.
Keep listening
Recent conversations

Why Pure Software Is Uninvestable: The SpaceX & YC Way to Build Hardware | Jamie Gull DD #12

Inside Deep Tech Investing: A Masterclass From Both Sides of the Table | Graeme Harrison
