Tweet

Tivadar Danka

16 Apr, 16 tweets, 4 min read

In the last 24 hours, more than 400 of you decided to follow me. Thank you, I am honored!

As you probably know, I love explaining complex machine learning concepts simply. I have collected some of my past threads for you to make sure you don't miss out on them.

Enjoy!

https://twitter.com/TivadarDanka/status/1359876189943382017

1. What is expected value?

https://twitter.com/TivadarDanka/status/1359876189943382017

https://twitter.com/TivadarDanka/status/1360237067826065411

2. What is entropy?

https://twitter.com/TivadarDanka/status/1360237067826065411

https://twitter.com/TivadarDanka/status/1361332789749104640

3. Why is matrix multiplication defined the way it is?

https://twitter.com/TivadarDanka/status/1361332789749104640

https://twitter.com/TivadarDanka/status/1361699203790106625

4. Where does the mean square error come from? An explanation from the Bayesian viewpoint.

https://twitter.com/TivadarDanka/status/1361699203790106625

https://twitter.com/TivadarDanka/status/1362422719049007107

5. The law of large numbers: clearing out one of the most common misconceptions about mathematics.

https://twitter.com/TivadarDanka/status/1362422719049007107

https://twitter.com/TivadarDanka/status/1363875041985908742

6. The Bayes formula unraveled.

https://twitter.com/TivadarDanka/status/1363875041985908742

https://twitter.com/TivadarDanka/status/1364226607020269574

7. The geometric and physics interpretation of differentiation.

https://twitter.com/TivadarDanka/status/1364226607020269574

https://twitter.com/TivadarDanka/status/1364582701504860161

8. What are the conditional probabilities?

https://twitter.com/TivadarDanka/status/1364582701504860161

https://twitter.com/TivadarDanka/status/1365319054425288709

9. Why are neural networks so effective? The universal approximation theorem, explained.

https://twitter.com/TivadarDanka/status/1365319054425288709

https://twitter.com/TivadarDanka/status/1367118665301295110

10. The best-kept secret of linear algebra: graph representation of matrices.

https://twitter.com/TivadarDanka/status/1367118665301295110

https://twitter.com/TivadarDanka/status/1375827708471566336

11. Vanishing gradients: the issues with Signoid.

https://twitter.com/TivadarDanka/status/1375827708471566336

https://twitter.com/TivadarDanka/status/1377261906461929473

12. What is gradient descent? An explanation with hill-climbing.

https://twitter.com/TivadarDanka/status/1377261906461929473

https://twitter.com/TivadarDanka/status/1379799882677022720

13. The central limit theorem: how scaled averages of distributions converge to Gaussian.

https://twitter.com/TivadarDanka/status/1379799882677022720

https://twitter.com/TivadarDanka/status/1381967666768846871

14. What is convolution? A simple interpretation with probability theory.

https://twitter.com/TivadarDanka/status/1381967666768846871

https://twitter.com/TivadarDanka/status/1382688415737593856

15. What does the inner product have to do with similarity? A geometric explanation.

https://twitter.com/TivadarDanka/status/1382688415737593856

• • •

Missing some Tweet in this thread? You can try to force a refresh

This Thread may be Removed Anytime!

Twitter may remove this content at anytime! Save it as PDF for later use!

More from @TivadarDanka

Tivadar Danka

@TivadarDanka

15 Apr

In machine learning, the inner product (or dot product) of vectors is often used to measure similarity.

However, the formula is far from revealing. What does the sum of coordinate products have to do with similarity?

There is a very simple geometric explanation!

🧵 👇🏽

There are two key things to observe.

First, the inner product is linear in both variables. This property is called bilinearity.

Second, is that the inner product is zero if the vectors are orthogonal.

Read 9 tweets

Tivadar Danka

@TivadarDanka

13 Apr

Convolution is not the easiest operation to understand: it involves functions, sums, and two moving parts.

However, there is an illuminating explanation — with probability theory!

There is a whole new aspect of convolution that you (probably) haven't seen before.

🧵 👇🏽

In machine learning, convolutions are most often applied for images, but to make our job easier, we shall take a step back and go to one dimension.

There, convolution is defined as below.

Now, let's forget about these formulas for a while, and talk about a simple probability distribution: we toss two 6-sided dices and study the resulting values.

To formalize the problem, let 𝑋 and 𝑌 be two random variables, describing the outcome of the first and second toss.

Read 9 tweets

Tivadar Danka

@TivadarDanka

8 Apr

One of my favorite convolutional network architectures is the U-Net.

It solves a hard problem in such an elegant way that it became one of the most performant and popular choices for semantic segmentation tasks.

How does it work?

🧵 👇🏽

Let's quickly recap what semantic segmentation is: a common computer vision task, where we want to classify which class each pixel belongs to.

Because we want to provide a prediction on a pixel level, this task is much harder than classification.

Since the absolutely classic paper Fully Convolutional Networks for Semantic Segmentation by Jonathan Long, Evan Shelhamer, and Trevor Darrell, fully end-to-end autoencoder architectures were most commonly used for this.

(Image source: paper above, arxiv.org/abs/1411.4038v2)

Read 10 tweets

Tivadar Danka

@TivadarDanka

7 Apr

There is a common misconception that all probability distributions are like a Gaussian.

Often, the reasoning involves the Central Limit Theorem.

This is not exactly right: they resemble Gaussian only from a certain perspective.

🧵 👇🏽

Let's state the CLT first. If we have 𝑋₁, 𝑋₂, ..., 𝑋ₙ independent and identically distributed random variables, their scaled sum is a Gaussian distribution in the limit.

The surprising thing here is the limit is independent of the variables' distribution.

Note that the random variables undergo a significant transformation: averaging and scaling with the mean, the variance, and √𝑛.

(The scaling transformation is the "certain perspective" I mentioned in the first tweet.)

Read 12 tweets

Tivadar Danka

@TivadarDanka

1 Apr

Gradient descent sounds good on paper, but there is a big issue in practice.

For complex functions like training losses for neural networks, calculating the gradient is computationally very expensive.

What makes it possible? For one, stochastic gradient descent!

🧵 👇🏽

When you have a lot of data, calculating the gradient of the loss involves the computation of a large sum.

Think about it: if 𝑥ᵢ denotes the data and 𝑤 denotes the weights, the loss function takes the form below.

Not only do we have to add a bunch of numbers together, but we have to find the gradient of each loss term.

For example, if the model contains 10 000 parameters and 1 000 000 data points, we need to compute 10 000 x 1 000 000 = 10¹⁰ derivatives.

This can take a LOT of time.

Read 7 tweets

Tivadar Danka

@TivadarDanka

31 Mar

Gradient descent has a really simple and intuitive explanation.

The algorithm is easy to understand once you realize that it is basically hill climbing with a really simple strategy.

Let's see how it works!

🧵 👇🏽

For functions of one variable, the gradient is simply the derivative of the function.

The derivative expresses the slope of the function's tangent plane, but it can also be viewed as a one-dimensional vector!

When the function is increasing, the derivative is positive. When decreasing, it is negative.

Translating this to the language of vectors, it means that the "gradient" points to the direction of the increase!

This is the key to understand gradient descent.

Read 9 tweets

Support us! We are indie developers!

This site is made by just two indie developers on a laptop doing marketing, support and development! Read more about the story.

Become a Premium Member ($3/month or $30/year) and get exclusive features!

Become Premium

Too expensive? Make a small donation by buying us coffee ($5) or help with server cost ($10)

Donate via Paypal Become our Patreon

Thank you for your support!

Share this page!

Tivadar Danka

Try unrolling a thread yourself!

More from @TivadarDanka

Tivadar Danka

Tivadar Danka

Tivadar Danka

Tivadar Danka

Tivadar Danka

Tivadar Danka

Did Thread Reader help you today?

Like this author's thread?