All transcripts and search

Local Minima - The Answer to Why We Are Seeking Minima

Uncorrected Scribe v2 machine transcript. Neutral speaker labels are specific to this recording. Review flags: 0.

TXT · SRT · VTT · Untouched JSON · Review flags

Source: youtube/20230309-oZoXwfHdLAw-Local Minima - The Answer to Why We Are Seeking Minima.mkv
SHA-256: 7642c564c22a22c7681433d7eb6311b78b61050e23135025cf5626d606ba2c0d

00:00:09.280 Speaker 1 Today we're gonna talk about one of my favorite niche topics. We're gonna talk about

00:00:16.200 Speaker 1 local minima. This is one of the concepts for which the channel itself is named.

00:00:22.280 Speaker 1 We're gonna talk about how local and global minima sort of matter to overall

00:00:27.940 Speaker 1 society, how you can think of them as why you get stuck in life, how to

00:00:33.780 Speaker 1 move on and grow as a person, how to improve things, how to learn better, and how to

00:00:39.660 Speaker 1 finally solve some of the problems you've been stuck on for a very long time.

00:00:43.520 Speaker 1 [clears throat] So first, I'm gonna start with a prediction, and then we're

00:00:49.320 Speaker 1 gonna back up into it. So here's the prediction. Reality will be

00:00:55.400 Speaker 1 seeded to the people who still know how to do things. What does

00:01:01.280 Speaker 1 that mean exactly? Right. Let's back up into what that prediction

00:01:07.000 Speaker 1 means. This is a prediction for something that will happen in the future. So I'll

00:01:10.980 Speaker 1 say it again. Reality will be seeded to the people who still know how

00:01:16.920 Speaker 1 to do things.

00:01:22.880 Speaker 1 For those of you who understand trading, you know a little bit about derivatives,

00:01:27.520 Speaker 1 and you know about paper value versus real value. What's the difference? Okay.

00:01:33.720 Speaker 1 So without diving too much into how currency works, let's say

00:01:39.640 Speaker 1 you have a piece of gold, a gold bar. Okay? A one-ounce

00:01:45.600 Speaker 1 gold bar. That is worth, you know, some amount of money or some amount of goods in

00:01:50.300 Speaker 1 the real world to other people who are willing to trade you that gold for something

00:01:54.280 Speaker 1 else. Okay? Now imagine I have an IOU. It's a piece of paper that says, "You can

00:02:00.140 Speaker 1 turn in this piece of paper for one ounce of gold at any point in time." That is

00:02:05.700 Speaker 1 paper value. It's not real value. Now, you can call your option

00:02:12.040 Speaker 1 on the real value anytime you want. You just go turn it in at the bank, and you can

00:02:15.320 Speaker 1 get your gold, right? It's almost as good. Now, those paper value vouchers are

00:02:21.320 Speaker 1 pretty awesome because you can trade them on exchanges. You can create all sorts of

00:02:25.100 Speaker 1 interesting derivatives. You can have leverage. You can do all sorts of crazy stuff

00:02:29.420 Speaker 1 with them.

00:02:31.380 Speaker 1 Here's where humanity gets into trouble.

00:02:34.940 Speaker 1 You can continue to print paper value in excess of real value. I can make

00:02:41.120 Speaker 1 100,000 vouchers that say, "You can turn in this voucher for a piece of gold."

00:02:47.100 Speaker 1 However, there may only be 1,000 pieces of gold. If people figure this out,

00:02:53.520 Speaker 1 they will quickly have a run on the gold. This is actually what a run on a bank is,

00:02:58.220 Speaker 1 is there is not enough money for you to go get it. It says you have this much in

00:03:02.480 Speaker 1 your bank account, but they actually don't have that much. Now, there's all sorts of

00:03:06.720 Speaker 1 layers of protections, but ultimately, that's how a lot of our economy works. The

00:03:11.940 Speaker 1 paper value is much larger than the real value, and under good times

00:03:17.780 Speaker 1 when nothing is under stress, that's not a big deal. Who cares, right? If paper

00:03:22.600 Speaker 1 value exceeds real value, that's only a problem if everybody tries to turn in their

00:03:26.220 Speaker 1 IOUs at the same time. It's a game of musical chairs, right? And that's the best

00:03:31.520 Speaker 1 analogy I can think of. There's not enough chairs. It's fine as long as the music is

00:03:35.660 Speaker 1 playing. But when the music stops, it's not so good.

00:03:40.800 Speaker 1 So now that you kind of understand the difference between paper and real value, I

00:03:45.720 Speaker 1 have a theory that... Well, this one is pretty well shared by everybody. Our economy

00:03:50.520 Speaker 1 is very over-financialized, meaning we made up a bunch of stuff on paper

00:03:57.220 Speaker 1 that isn't quite backed by the same amount of real value, the ability to

00:04:02.320 Speaker 1 produce goods, the real goods themselves in storage, the ability to do useful

00:04:08.260 Speaker 1 stuff in the real world. This is gonna be a running theme through all of Seeking

00:04:12.440 Speaker 1 Minima. There is no replacement for being good at doing

00:04:18.459 Speaker 1 stuff that matters to people. There's no replacement for that. No investment, no

00:04:24.440 Speaker 1 idea, nothing is as good as being able to do something in the real world

00:04:30.800 Speaker 1 that is useful to another person. All right.

00:04:35.920 Speaker 1 So besides over-financialized economy, we also have an

00:04:41.360 Speaker 1 over-specialized economy. I call this the tall tower problem.

00:04:48.100 Speaker 1 So in finance, there's layers upon layers of derivatives, and you can see when this

00:04:53.520 Speaker 1 becomes an option, like, mm, back in 2007 when there was the subprime mortgage

00:04:59.020 Speaker 1 lending crisis. Basically, we just had stacks and stacks of derivatives. They were

00:05:03.200 Speaker 1 all being shuffled and bundled around, and underneath it all were people's home

00:05:08.140 Speaker 1 mortgages. But there were so many layers of derivatives, nobody even knew who owned

00:05:13.200 Speaker 1 what anymore. Right? It was all paper value on paper value on paper value, many

00:05:18.980 Speaker 1 orders away from a real thing, somebody's house. We have something kind

00:05:24.880 Speaker 1 of like that happening with specialization.

00:05:29.080 Speaker 1 So think of a very tall tower. All right? Let's say you have a specialized num- a

00:05:34.020 Speaker 1 certain number of blocks. Each block represents a skill, a

00:05:39.360 Speaker 1 capacity you can learn. Right? It's your ability to do something,

00:05:45.860 Speaker 1 and you can choose to build a tower that is tall. Right? You can be... That is

00:05:51.640 Speaker 1 narrow, deep expertise. You can be an incredible microbiologist on a

00:05:57.720 Speaker 1 very specific strain of bacteria. And you can be very good at that thing. And that

00:06:03.412 Speaker 1 can be incredibly useful to the world as long as that's what the world needs. Or you

00:06:08.012 Speaker 1 can also know how to repair a car. You can also know how to fix your bike. You can

00:06:13.052 Speaker 1 also know how to set up your own solar array. Like there's a lot of other things you

00:06:18.152 Speaker 1 could be learning besides microbiology, a very specific kind, but you won't be as

00:06:22.432 Speaker 1 good at any of them, right? We all sort of face this trade-off. Do I go deep? Do I

00:06:26.532 Speaker 1 go broad? So we all have a certain number of skill blocks, finite cap on how much

00:06:31.852 Speaker 1 stuff we can learn.

00:06:35.032 Speaker 1 Now, what happens, here's, here's where these two things come into contact really,

00:06:41.212 Speaker 1 is the over-financialized economy really, really loves specialization because

00:06:46.132 Speaker 1 specialization is great. It ensures that you can be the best at the

00:06:51.552 Speaker 1 forty-first layer on your tall tower, right? So if you just go super deep and you

00:06:56.472 Speaker 1 just say, "Well, I'm gonna make the most money in a hyper-financialized economy,"

00:07:01.412 Speaker 1 the best way to do that is just to get really amazing at something super high value.

00:07:06.291 Speaker 1 These are often in technology. It's like engineering, medicine, sometimes

00:07:12.212 Speaker 1 law, or finance, finance itself, right? It's such an esoteric

00:07:18.152 Speaker 1 thing when you get down and actually think about it. Think about what your job is

00:07:21.772 Speaker 1 right now. What layer of a tall tower are you standing on, right? Are you the

00:07:29.252 Speaker 1 digital marketer for Facebook, for pet, for dog walking companies?

00:07:35.272 Speaker 1 Wow, that is very specific. Now, it makes sense to make a living in a

00:07:40.232 Speaker 1 hyper-financialized economy being very specialized. That's where the money is. You

00:07:44.352 Speaker 1 want to be able to distinguish yourself. If you just say, "Ah, I do everything,"

00:07:48.692 Speaker 1 nobody will hire you because they don't want somebody who does everything. They want

00:07:51.612 Speaker 1 somebody who is the best at the exact thing they want. That's cool.

00:07:56.492 Speaker 1 But the whole point of this is to say that there are downsides to the tall tower.

00:08:01.232 Speaker 1 Here's an example. Facebook goes away. Now who are you?

00:08:07.392 Speaker 1 You're the marketer for Facebook ads for dog walking companies. Sure. Okay. Maybe

00:08:12.312 Speaker 1 you can pivot a little bit to another platform. Certainly. Yeah. But what happens if

00:08:18.112 Speaker 1 the ad model for businesses changes on the internet? What if, I don't know,

00:08:23.812 Speaker 1 something like a generalized, generalized language model makes it so that people

00:08:28.312 Speaker 1 just aren't looking at very many ads anymore? Hmm. [laughs] Your

00:08:34.232 Speaker 1 economy, ad economy implodes. There's just like ninety percent less revenue overall

00:08:39.652 Speaker 1 in the future. You are sitting at the very top of a tall tower that has just been

00:08:44.912 Speaker 1 shaken, and twenty of those blocks fall off. Now what do you know?

00:08:51.052 Speaker 1 Nothing. You're useless. You don't know anything. You don't know how to fix your

00:08:55.872 Speaker 1 car. You don't know how to fix your bike. You don't know how to set up a solar

00:08:59.112 Speaker 1 array. You don't even know, like, how to pivot to something similar because you've

00:09:03.672 Speaker 1 gotten so good at one thing and only one thing.

00:09:07.732 Speaker 1 This is the trouble with tall towers, is that they're not very stable. I'm gonna

00:09:11.932 Speaker 1 extrapolate this even further and say that our entire society

00:09:18.092 Speaker 1 is like this. We have a very, very, very tall tower, and we just

00:09:24.112 Speaker 1 keep putting one more block on top of the other, training people to be some

00:09:29.172 Speaker 1 hyper-specialized thing, while really understanding very little of the fundamentals

00:09:33.792 Speaker 1 underneath it or anything adjacent to them. Now, this is certainly not a

00:09:39.672 Speaker 1 rule, but people who are sort of well-rounded and good at multiple things are

00:09:44.172 Speaker 1 definitely the exception. And this is not to shame people who have become

00:09:48.952 Speaker 1 specialized. It made sense in our economy, in our world. But if we're ever

00:09:55.052 Speaker 1 shaken, as we sometimes are, um, we'll say with some regularity maybe, fi- every

00:10:00.912 Speaker 1 five to ten years, let's say all of society is shaken by something. Are you gonna be

00:10:05.852 Speaker 1 one of those blocks that gets shaken out? I think that we're in a very fragile place

00:10:11.912 Speaker 1 overall in terms of human civilization because our tower is so tall

00:10:17.991 Speaker 1 and it's not robust.

00:10:21.032 Speaker 1 So now we're gonna talk about how this fits into minimas, specifically local

00:10:26.532 Speaker 1 minimas. So I think a lot about how machine learning models actually learn, and

00:10:32.552 Speaker 1 I think there's so many analogies that you can draw to real life. When you look at

00:10:37.012 Speaker 1 how a non-human agent learns, there's just like philosophical wonder in

00:10:42.832 Speaker 1 that. I don't know. I, I can draw so many con- so many interesting parallels to real

00:10:46.612 Speaker 1 life. So when a machine learning model starts to learn, in this case, I'm talking

00:10:51.912 Speaker 1 about deep neural networks. When they start to learn something, we'll say

00:10:55.592 Speaker 1 recognizing objects in photos, they will start with a pretty broad,

00:11:01.312 Speaker 1 messy approach, right? Let's just try stuff. Kind of like a baby, just trying random

00:11:06.112 Speaker 1 things that makes no sense. And that's-- so it starts out broad, and then as soon as

00:11:11.152 Speaker 1 it starts to find a winning strategy, it will start to go deep, right? Start to

00:11:14.892 Speaker 1 build up that tower. It'll double down on this technique. So what you want as the

00:11:20.892 Speaker 1 machine learning engineer is you want a model that learns lots of different ways to

00:11:25.552 Speaker 1 detect things. It has a broad base to its tower, but it does get very good

00:11:32.072 Speaker 1 as it-- at the actual task, right? You want to actually be able to tell a cat

00:11:38.272 Speaker 1 from a bicycle or a cat from a cheetah. Chelling- telling a cat from a cheetah is

00:11:43.452 Speaker 1 very difficult. So you want it to be great, a deep specialist to handle those

00:11:48.932 Speaker 1 difficult edge cases, but you want it to be generalized as well. It can also tell a

00:11:54.032 Speaker 1 cat from a bus, which sounds silly to us to say that. But if you have a very narrow

00:11:59.732 Speaker 1 model and it's not well generalized, that's exactly what it can't do. [laughs] It

00:12:04.252 Speaker 1 doesn't know the difference between a cat and a bus. It, it doesn't understand even

00:12:07.172 Speaker 1 how to conceive of these problems.

00:12:10.942 Speaker 1 So this problem is called overfitting. When you train-- when a machine

00:12:16.822 Speaker 1 learning model doubles down on one strategy too much, it goes too deep. And in this

00:12:21.282 Speaker 1 case, like our tall tower problem, this is a tower that is too tall and is flimsy.

00:12:28.362 Speaker 1 That's called overfitting. It basically tries to memorize the data. It finds a

00:12:33.482 Speaker 1 technique that works well, and it only does this one thing, no matter what.

00:12:40.042 Speaker 1 This is what it's like to be s-stuck in a local minima. So if you take a graph,

00:12:47.642 Speaker 1 we'll say every point on the graph represents how good the model is. It's literally

00:12:52.982 Speaker 1 its score for how good it is at its specific problem, right? So higher points on the

00:12:58.522 Speaker 1 graph, in this case, represent bad scores. Lower points on the graph

00:13:04.342 Speaker 1 represent good scores. Now, if we have the x-axis as time, then you will

00:13:09.942 Speaker 1 see that over time, you know, it's gonna go up and down, right? Typically, it starts

00:13:15.822 Speaker 1 bad, and then it gets better, right? Up and down, up and down. Starts to learn a

00:13:21.202 Speaker 1 little bit. It messes up a little bit, tries a new technique, and it learns over

00:13:24.922 Speaker 1 time. You expect over time, it's gonna go down, meaning it's learning better. Lower

00:13:30.982 Speaker 1 points means it's better over time. So you can think of it kind of like a hilly

00:13:36.402 Speaker 1 terrain, right? It's got peaks and valleys, kind of like a sine wave.

00:13:43.322 Speaker 1 The valleys are local minima, right? It is the lowest point on a

00:13:49.162 Speaker 1 curve. Machine learning models can get stuck at this low point on the curve.

00:13:55.342 Speaker 1 They can never leave. They can't leave because they found a strategy that works

00:14:00.562 Speaker 1 well, but they can't find a strategy that is

00:14:06.742 Speaker 1 better overall without first making their own score worse. In

00:14:12.662 Speaker 1 order to actually get better overall, it actually has to get worse for a little bit.

00:14:19.002 Speaker 1 So let's draw an analogy to real life. Let's say you are at block number forty on

00:14:24.082 Speaker 1 your tall tower. You are the Facebook marketer [chuckles] for a dog walking

00:14:28.822 Speaker 1 companies. Well, if you want to get better overall, you--

00:14:35.122 Speaker 1 as in more robust to tower shaking, and you wanna be better at overall life, less

00:14:40.142 Speaker 1 fragile, you would probably need to learn an adjacent skill. Maybe you learn, like,

00:14:45.382 Speaker 1 generalized copywriting, or maybe you also-- or just learn a separate tower, right?

00:14:50.962 Speaker 1 You also learn photography. Okay? Something like this. It's sort of relevant to what

00:14:55.982 Speaker 1 you're doing, but not exactly the same thing. So in this case, when you start

00:15:01.582 Speaker 1 learning photography, you're gonna suck at it. You're gonna be bad. It's literally

00:15:05.882 Speaker 1 going to make you worse overall. You're gonna make-- be made worse overall because

00:15:10.562 Speaker 1 you can't focus on the one thing you're great at, and you have to start out bad. A

00:15:15.382 Speaker 1 lot of people get stuck here. They get to a comfort zone in life. They get

00:15:21.322 Speaker 1 pretty good at one thing, and then they stop. They get stuck in a local minima, just

00:15:26.982 Speaker 1 like a machine learning model. They can't really branch out and try new things. They

00:15:31.622 Speaker 1 can't really get into a new career, a new job. They can't really go back to school

00:15:36.962 Speaker 1 because of the, the perceived risk. The penalty for trying to be something new, more

00:15:42.762 Speaker 1 flexible, more robust is too high. They can't get over the hump

00:15:48.862 Speaker 1 to get to a better place. Even if there were a lower valley, remember, low points

00:15:53.362 Speaker 1 represent better overall fitness score at the problem you're trying to solve, they

00:15:58.222 Speaker 1 can't get to this lower point because they keep getting stuck. They can't quite

00:16:02.942 Speaker 1 build the momentum to get over the hill, so they stay where they are in a local

00:16:06.902 Speaker 1 minima.

00:16:09.522 Speaker 1 Again, I think this is where humanity is. We are [chuckles] currently stuck in a

00:16:14.742 Speaker 1 local minima. We can't get to the global minima. We can't even get to lower local

00:16:19.142 Speaker 1 minima because we're comfortable, because we figured things out.

00:16:25.082 Speaker 1 So how do you fix it? Again, we can draw analogies to machine learning. In machine

00:16:30.901 Speaker 1 learning, you have a lot of different strategies to sort of

00:16:34.522 Speaker 1 get models to try new things, prevent them from over-optimizing. Okay, cool, you

00:16:39.402 Speaker 1 found one strategy that works, but basically, you force it to be flexible. You force

00:16:44.702 Speaker 1 it to be unable to use its one strategy that it finds, so it can't be a one-trick

00:16:49.402 Speaker 1 pony. So what do you do? You do things like dropout. This means literally removing

00:16:54.822 Speaker 1 parts of its neurons, right? So basically, parts of the network, you turn them off.

00:17:00.382 Speaker 1 This is like making you selectively forget so that you have to find a more general

00:17:04.442 Speaker 1 solution, right? Some of the clues are not always there. Some of the ways you did

00:17:09.202 Speaker 1 things literally removes parts of the network. We can't really do that with humans.

00:17:13.302 Speaker 1 Um, here's another one, data augmentation. Right. So th-- for the picture example,

00:17:18.702 Speaker 1 instead of just sending the same pictures of cats and buses, sometimes you flip

00:17:23.002 Speaker 1 them, sometimes you make them kind of blurry, sometimes you zoom in on them and crop

00:17:28.062 Speaker 1 them so that it can't just memorize things, and it can't always look for a perfect

00:17:33.042 Speaker 1 picture of a cat right in the center of the frame. It has to kind of learn, what do

00:17:37.422 Speaker 1 you do when it's off-center a little bit? What do you do when it's kind of blurry?

00:17:41.402 Speaker 1 So it has to learn other approaches to win.

00:17:46.742 Speaker 1 So there are other things you can do. Noise injection. You can do loss function

00:17:51.842 Speaker 1 penalties. Um, so this is a way of saying l-- uh, nonlinear loss function penalties.

00:17:57.202 Speaker 1 This is a way of saying, if you're a little bit wrong, that's okay. If you're really

00:18:01.622 Speaker 1 wrong, we're gonna penalize you, not just, like, the same amount more. If you're

00:18:06.702 Speaker 1 fifty percent wrong, you don't get fifty percent penalty. You get five hundred

00:18:11.882 Speaker 1 percent penalty, right? It goes exponential. So if you're really bad, you're way off

00:18:17.462 Speaker 1 on something, we're gonna penalize you severely. We can do things like that in real

00:18:23.032 Speaker 1 life too. This is like saying, instead of I'm already on the 41st block

00:18:29.072 Speaker 1 with my day job. Do I really need to be at the 42nd block by

00:18:34.832 Speaker 1 learning this other niche tool to be a better Facebook dog walking

00:18:40.532 Speaker 1 marketer? Or maybe I could learn how

00:18:46.512 Speaker 1 to seal my radiator in my car, right? Even though it's not optimal.

00:18:52.791 Speaker 1 Maybe I could do something adjacent, right? Like learning photography or learning

00:18:58.152 Speaker 1 something else. Become more robust. Not because you have to, but just in case the

00:19:04.112 Speaker 1 tower gets shaken, right? And really, that's all you need to do

00:19:10.612 Speaker 1 is deliberately shake yourself. And that's all humanity needs is

00:19:16.212 Speaker 1 sometimes we need to be shaken. Because if we're not, we just keep building the

00:19:21.192 Speaker 1 tower taller. We keep being stuck in local minima.

00:19:26.992 Speaker 1 Hyper-optimizing for a tiny, tiny percentage point of a gain with our 10th layer of

00:19:33.252 Speaker 1 financial derivative. This is how you get content that isn't content.

00:19:39.732 Speaker 1 It's vapid. There's nothing there.

00:19:44.172 Speaker 1 I think I've said that content is becoming double speak. What does it even mean? You

00:19:50.032 Speaker 1 can't just pump this stuff out. What are we doing? It's very clear to me that we're

00:19:56.012 Speaker 1 not just over-financialized, we're over-specialized. Too many layers of derivatives

00:20:01.092 Speaker 1 at every level. And the answer is shake it. Because we're stuck in a local

00:20:07.112 Speaker 1 minima. And if we ever want to get to the global minima of our society, we should

00:20:13.092 Speaker 1 thoughtfully shake ourselves. We should thoughtfully start becoming more robust and

00:20:18.932 Speaker 1 prepared and well-rounded where we can. So that when the tower does fall,

00:20:24.712 Speaker 1 because it will fall,

00:20:27.832 Speaker 1 hopefully it doesn't fall all the way down to the first block, if you catch my

00:20:31.772 Speaker 1 drift. So that's how I think about local minima and why I think they're a useful

00:20:37.692 Speaker 1 concept. We get stuck because we find something that works well,

00:20:44.012 Speaker 1 right? This is sort of analogous to

00:20:49.412 Speaker 1 being happy, being good instead of great, you know, from the book Good to Great.

00:20:54.192 Speaker 1 Fantastic concept. Very similar. You'll find analogies all over the place. But the

00:20:58.432 Speaker 1 way I like to think of it is local minima. Because I can just imagine myself being

00:21:01.772 Speaker 1 on a bike, like stuck at the bottom of a hill in a valley between two hills. And

00:21:07.612 Speaker 1 even though I know that there's a lower point farther away, I can't get there

00:21:12.792 Speaker 1 because I'm not willing to climb this hill on my bike. So that is my advice.

00:21:19.672 Speaker 1 Begin to notice when you're stuck in a local minima. When you're stuck,

00:21:24.752 Speaker 1 you need to start climbing. You need to see what's over the next hill. Because

00:21:29.892 Speaker 1 chances are, it's a lower point. It's a better local minima.