Local Minima - The Answer to Why We Are Seeking Minima
Seeking Minima
This original recording is preserved in MKV format. Download the saved file for a compatible video player, or watch it at its original source.
Read transcript
Machine transcription; unreviewed and may contain errors.
Local Minima - The Answer to Why We Are Seeking Minima Source: youtube/20230309-oZoXwfHdLAw-Local Minima - The Answer to Why We Are Seeking Minima.mkv SHA-256: 7642c564c22a22c7681433d7eb6311b78b61050e23135025cf5626d606ba2c0d Model: scribe_v2 | Transcribed: 2026-10-10T18:25:00.440490+00:00 Machine transcript — uncorrected. Speaker labels are local to this recording and do not identify people. [00:00:09.280] Speaker 1: Today we're gonna talk about one of my favorite niche topics. We're gonna talk about local minima. This is one of the concepts for which the channel itself is named. We're gonna talk about how local and global minima sort of matter to overall society, how you can think of them as why you get stuck in life, how to move on and grow as a person, how to improve things, how to learn better, and how to finally solve some of the problems you've been stuck on for a very long time. [00:00:43.520] Speaker 1: [clears throat] So first, I'm gonna start with a prediction, and then we're gonna back up into it. So here's the prediction. Reality will be seeded to the people who still know how to do things. What does that mean exactly? Right. Let's back up into what that prediction means. This is a prediction for something that will happen in the future. So I'll say it again. Reality will be seeded to the people who still know how to do things. [00:01:22.880] Speaker 1: For those of you who understand trading, you know a little bit about derivatives, and you know about paper value versus real value. What's the difference? Okay. So without diving too much into how currency works, let's say you have a piece of gold, a gold bar. Okay? A one-ounce gold bar. That is worth, you know, some amount of money or some amount of goods in the real world to other people who are willing to trade you that gold for something else. Okay? Now imagine I have an [00:01:57.540] Speaker 1: IOU. It's a piece of paper that says, "You can turn in this piece of paper for one ounce of gold at any point in time." That is paper value. It's not real value. Now, you can call your option on the real value anytime you want. You just go turn it in at the bank, and you can get your gold, right? It's almost as good. Now, those paper value vouchers are pretty awesome because you can trade them on exchanges. You can create all sorts of interesting derivatives. You can have leverage. You can do all sorts of crazy stuff with them. Here's where humanity [00:02:32.400] Speaker 1: gets into trouble. You can continue to print paper value in excess of real value. I can make 100,000 vouchers that say, "You can turn in this voucher for a piece of gold." However, there may only be 1,000 pieces of gold. If people figure this out, they will quickly have a run on the gold. This is actually what a run on a bank is, is there is not enough money for you to go get it. It says you have this much in your bank account, but they actually don't have that much. Now, there's all sorts of layers of [00:03:07.200] Speaker 1: protections, but ultimately, that's how a lot of our economy works. The paper value is much larger than the real value, and under good times when nothing is under stress, that's not a big deal. Who cares, right? If paper value exceeds real value, that's only a problem if everybody tries to turn in their IOUs at the same time. It's a game of musical chairs, right? And that's the best analogy I can think of. There's not enough chairs. It's fine as long as the music is playing. But when the music stops, it's not so good. So [00:03:42.300] Speaker 1: now that you kind of understand the difference between paper and real value, I have a theory that... Well, this one is pretty well shared by everybody. Our economy is very over-financialized, meaning we made up a bunch of stuff on paper that isn't quite backed by the same amount of real value, the ability to produce goods, the real goods themselves in storage, the ability to do useful stuff in the real world. This is gonna be a running theme through all of Seeking Minima. There is no replacement for being [00:04:17.079] Speaker 1: good at doing stuff that matters to people. There's no replacement for that. No investment, no idea, nothing is as good as being able to do something in the real world that is useful to another person. All right. So besides over-financialized economy, we also have an over-specialized economy. I call this the tall tower problem. So in finance, there's layers upon layers of derivatives, [00:04:52.580] Speaker 1: and you can see when this becomes an option, like, mm, back in 2007 when there was the subprime mortgage lending crisis. Basically, we just had stacks and stacks of derivatives. They were all being shuffled and bundled around, and underneath it all were people's home mortgages. But there were so many layers of derivatives, nobody even knew who owned what anymore. Right? It was all paper value on paper value on paper value, many orders away from a real thing, somebody's house. We have something kind of like that happening with specialization. [00:05:29.080] Speaker 1: So think of a very tall tower. All right? Let's say you have a specialized num- a certain number of blocks. Each block represents a skill, a capacity you can learn. Right? It's your ability to do something, and you can choose to build a tower that is tall. Right? You can be... That is narrow, deep expertise. You can be an incredible microbiologist on a very specific strain of bacteria. And you can be very good at that thing. And that can be incredibly [00:06:04.112] Speaker 1: useful to the world as long as that's what the world needs. Or you can also know how to repair a car. You can also know how to fix your bike. You can also know how to set up your own solar array. Like there's a lot of other things you could be learning besides microbiology, a very specific kind, but you won't be as good at any of them, right? We all sort of face this trade-off. Do I go deep? Do I go broad? So we all have a certain number of skill blocks, finite cap on how much stuff we can learn. Now, what happens, here's, here's where these [00:06:39.052] Speaker 1: two things come into contact really, is the over-financialized economy really, really loves specialization because specialization is great. It ensures that you can be the best at the forty-first layer on your tall tower, right? So if you just go super deep and you just say, "Well, I'm gonna make the most money in a hyper-financialized economy," the best way to do that is just to get really amazing at something super high value. These are often in technology. It's like engineering, medicine, sometimes law, or [00:07:13.952] Speaker 1: finance, finance itself, right? It's such an esoteric thing when you get down and actually think about it. Think about what your job is right now. What layer of a tall tower are you standing on, right? Are you the digital marketer for Facebook, for pet, for dog walking companies? Wow, that is very specific. Now, it makes sense to make a living in a hyper-financialized economy being very specialized. That's where the money is. You want to be able to distinguish yourself. If you just say, "Ah, I do everything," [00:07:48.692] Speaker 1: nobody will hire you because they don't want somebody who does everything. They want somebody who is the best at the exact thing they want. That's cool. But the whole point of this is to say that there are downsides to the tall tower. Here's an example. Facebook goes away. Now who are you? You're the marketer for Facebook ads for dog walking companies. Sure. Okay. Maybe you can pivot a little bit to another platform. Certainly. Yeah. But what happens if the ad model for businesses changes on the internet? What if, I don't know, [00:08:23.812] Speaker 1: something like a generalized, generalized language model makes it so that people just aren't looking at very many ads anymore? Hmm. [laughs] Your economy, ad economy implodes. There's just like ninety percent less revenue overall in the future. You are sitting at the very top of a tall tower that has just been shaken, and twenty of those blocks fall off. Now what do you know? Nothing. You're useless. You don't know anything. You don't know how to fix your car. You don't know how to fix your bike. You don't know how to set up a [00:08:58.792] Speaker 1: solar array. You don't even know, like, how to pivot to something similar because you've gotten so good at one thing and only one thing. This is the trouble with tall towers, is that they're not very stable. I'm gonna extrapolate this even further and say that our entire society is like this. We have a very, very, very tall tower, and we just keep putting one more block on top of the other, training people to be some hyper-specialized thing, while really understanding very little of the fundamentals [00:09:33.792] Speaker 1: underneath it or anything adjacent to them. Now, this is certainly not a rule, but people who are sort of well-rounded and good at multiple things are definitely the exception. And this is not to shame people who have become specialized. It made sense in our economy, in our world. But if we're ever shaken, as we sometimes are, um, we'll say with some regularity maybe, fi- every five to ten years, let's say all of society is shaken by something. Are you gonna be one of those blocks that gets shaken out? I [00:10:08.792] Speaker 1: think that we're in a very fragile place overall in terms of human civilization because our tower is so tall and it's not robust. So now we're gonna talk about how this fits into minimas, specifically local minimas. So I think a lot about how machine learning models actually learn, and I think there's so many analogies that you can draw to real life. When you look at how a non-human agent learns, there's just like philosophical wonder in that. I don't know. [00:10:43.792] Speaker 1: I, I can draw so many con- so many interesting parallels to real life. So when a machine learning model starts to learn, in this case, I'm talking about deep neural networks. When they start to learn something, we'll say recognizing objects in photos, they will start with a pretty broad, messy approach, right? Let's just try stuff. Kind of like a baby, just trying random things that makes no sense. And that's-- so it starts out broad, and then as soon as it starts to find a winning strategy, it will start to go deep, right? Start to build up that tower. It'll double down on this technique. [00:11:18.832] Speaker 1: So what you want as the machine learning engineer is you want a model that learns lots of different ways to detect things. It has a broad base to its tower, but it does get very good as it-- at the actual task, right? You want to actually be able to tell a cat from a bicycle or a cat from a cheetah. Chelling- telling a cat from a cheetah is very difficult. So you want it to be great, a deep specialist to handle those difficult edge cases, but you want it to be generalized as well. It can also tell [00:11:53.912] Speaker 1: a cat from a bus, which sounds silly to us to say that. But if you have a very narrow model and it's not well generalized, that's exactly what it can't do. [laughs] It doesn't know the difference between a cat and a bus. It, it doesn't understand even how to conceive of these problems. So this problem is called overfitting. When you train-- when a machine learning model doubles down on one strategy too much, it goes too deep. And in this case, like our tall tower problem, this is a tower that is too tall and is flimsy. That's called [00:12:28.842] Speaker 1: overfitting. It basically tries to memorize the data. It finds a technique that works well, and it only does this one thing, no matter what. This is what it's like to be s-stuck in a local minima. So if you take a graph, we'll say every point on the graph represents how good the model is. It's literally its score for how good it is at its specific problem, right? So higher points on the graph, in this case, represent bad scores. Lower points on the [00:13:03.622] Speaker 1: graph represent good scores. Now, if we have the x-axis as time, then you will see that over time, you know, it's gonna go up and down, right? Typically, it starts bad, and then it gets better, right? Up and down, up and down. Starts to learn a little bit. It messes up a little bit, tries a new technique, and it learns over time. You expect over time, it's gonna go down, meaning it's learning better. Lower points means it's better over time. So you can think of it kind of like a hilly terrain, right? It's [00:13:38.502] Speaker 1: got peaks and valleys, kind of like a sine wave. The valleys are local minima, right? It is the lowest point on a curve. Machine learning models can get stuck at this low point on the curve. They can never leave. They can't leave because they found a strategy that works well, but they can't find a strategy that is better overall without first making their own score worse. In order to [00:14:13.302] Speaker 1: actually get better overall, it actually has to get worse for a little bit. So let's draw an analogy to real life. Let's say you are at block number forty on your tall tower. You are the Facebook marketer [chuckles] for a dog walking companies. Well, if you want to get better overall, you-- as in more robust to tower shaking, and you wanna be better at overall life, less fragile, you would probably need to learn an adjacent skill. Maybe you learn, like, generalized copywriting, or maybe you also-- or [00:14:48.282] Speaker 1: just learn a separate tower, right? You also learn photography. Okay? Something like this. It's sort of relevant to what you're doing, but not exactly the same thing. So in this case, when you start learning photography, you're gonna suck at it. You're gonna be bad. It's literally going to make you worse overall. You're gonna make-- be made worse overall because you can't focus on the one thing you're great at, and you have to start out bad. A lot of people get stuck here. They get to a comfort zone in life. They get pretty good at one thing, and then [00:15:23.202] Speaker 1: they stop. They get stuck in a local minima, just like a machine learning model. They can't really branch out and try new things. They can't really get into a new career, a new job. They can't really go back to school because of the, the perceived risk. The penalty for trying to be something new, more flexible, more robust is too high. They can't get over the hump to get to a better place. Even if there were a lower valley, remember, low points represent better overall fitness score at the problem you're trying to solve, they [00:15:58.222] Speaker 1: can't get to this lower point because they keep getting stuck. They can't quite build the momentum to get over the hill, so they stay where they are in a local minima. Again, I think this is where humanity is. We are [chuckles] currently stuck in a local minima. We can't get to the global minima. We can't even get to lower local minima because we're comfortable, because we figured things out. So how do you fix it? Again, we can draw analogies to machine learning. In machine learning, you have a lot of different strategies to sort of [00:16:34.522] Speaker 1: get models to try new things, prevent them from over-optimizing. Okay, cool, you found one strategy that works, but basically, you force it to be flexible. You force it to be unable to use its one strategy that it finds, so it can't be a one-trick pony. So what do you do? You do things like dropout. This means literally removing parts of its neurons, right? So basically, parts of the network, you turn them off. This is like making you selectively forget so that you have to find a more general solution, right? Some of the clues are not always there. Some of the ways you did things [00:17:09.682] Speaker 1: literally removes parts of the network. We can't really do that with humans. Um, here's another one, data augmentation. Right. So th-- for the picture example, instead of just sending the same pictures of cats and buses, sometimes you flip them, sometimes you make them kind of blurry, sometimes you zoom in on them and crop them so that it can't just memorize things, and it can't always look for a perfect picture of a cat right in the center of the frame. It has to kind of learn, what do you do when it's off-center a little bit? What do you do when it's kind of blurry? So it has to learn other approaches to [00:17:44.582] Speaker 1: win. So there are other things you can do. Noise injection. You can do loss function penalties. Um, so this is a way of saying l-- uh, nonlinear loss function penalties. This is a way of saying, if you're a little bit wrong, that's okay. If you're really wrong, we're gonna penalize you, not just, like, the same amount more. If you're fifty percent wrong, you don't get fifty percent penalty. You get five hundred percent penalty, right? It goes exponential. So if you're really bad, you're way off on something, we're gonna penalize you [00:18:19.102] Speaker 1: severely. We can do things like that in real life too. This is like saying, instead of I'm already on the 41st block with my day job. Do I really need to be at the 42nd block by learning this other niche tool to be a better Facebook dog walking marketer? Or maybe I could learn how to seal my radiator in my car, right? Even though it's not optimal. Maybe I could do [00:18:54.032] Speaker 1: something adjacent, right? Like learning photography or learning something else. Become more robust. Not because you have to, but just in case the tower gets shaken, right? And really, that's all you need to do is deliberately shake yourself. And that's all humanity needs is sometimes we need to be shaken. Because if we're not, we just keep building the tower taller. We keep being stuck in local minima. Hyper-optimizing for a [00:19:28.692] Speaker 1: tiny, tiny percentage point of a gain with our 10th layer of financial derivative. This is how you get content that isn't content. It's vapid. There's nothing there. I think I've said that content is becoming double speak. What does it even mean? You can't just pump this stuff out. What are we doing? It's very clear to me that we're not just over-financialized, we're over-specialized. Too many layers of derivatives at every level. And the answer is [00:20:04.512] Speaker 1: shake it. Because we're stuck in a local minima. And if we ever want to get to the global minima of our society, we should thoughtfully shake ourselves. We should thoughtfully start becoming more robust and prepared and well-rounded where we can. So that when the tower does fall, because it will fall, hopefully it doesn't fall all the way down to the first block, if you catch my drift. So that's how I think about local minima and why I think they're a useful concept. [00:20:39.492] Speaker 1: We get stuck because we find something that works well, right? This is sort of analogous to being happy, being good instead of great, you know, from the book Good to Great. Fantastic concept. Very similar. You'll find analogies all over the place. But the way I like to think of it is local minima. Because I can just imagine myself being on a bike, like stuck at the bottom of a hill in a valley between two hills. And even though I know that there's a lower point farther away, I can't get there because I'm not willing to [00:21:14.172] Speaker 1: climb this hill on my bike. So that is my advice. Begin to notice when you're stuck in a local minima. When you're stuck, you need to start climbing. You need to see what's over the next hill. Because chances are, it's a lower point. It's a better local minima.
Date basis: YouTube upload date. The saved file is a preservation copy, retained without re-encoding. Source records and checksums are available in Downloads.