← All posts

We Know More Than We Can Tell

There is a simple sentence from philosopher Michael Polanyi that I keep coming back to while learning more about AI:

“We can know more than we can tell.”

This idea, commonly associated with Polanyi's paradox, becomes surprisingly interesting when we think about how humans learn versus how we train machines.

Think about something very simple. You see a dog walking down the street and almost instantly know that it is a dog. Your brain doesn't consciously run through a checklist like: four legs — check, tail — check, fur — check, dog-shaped ears — check. You simply look at it and know.

But now suppose someone asks you to write down the exact rules you used to recognize that dog. Suddenly it becomes difficult. A cat also has four legs and fur. A wolf may look very similar to a dog. A Chihuahua and a Great Dane barely look alike, yet even a child recognizes both as dogs.

That is the fascinating part: we know what a dog is, but we cannot completely explain all the rules our brain uses to know it. We know more than we can tell.

So how would we teach a computer?

If we used traditional programming, we might start with something like this:

def is_dog(image):
    if has_four_legs(image) and has_fur(image) and has_tail(image):
        return True

    return False

This looks reasonable until we show it a cat. Or a wolf. Or a three-legged dog. Or just the face of a dog looking through a window.

Our rules quickly fall apart.

Machine learning takes a different approach. Instead of trying to explain every rule, we give the machine lots of examples:

🐶 → Dog
🐱 → Not Dog
🐕 → Dog
🐯 → Not Dog
🦮 → Dog

The important difference is simple:

Traditional programming:

Rules + Data  →  Answer


Machine Learning:

Data + Answers  →  Learned Rules

We are effectively saying: “I can't completely explain what makes something a dog, but here are thousands of dogs. Figure out the patterns.”

And that idea is much closer to Polanyi's paradox.

What is actually happening inside the neural network?

An artificial neuron can be simplified to something like:

z = (x1 * w1) + (x2 * w2) + bias

output = relu(z)

The x values are inputs and the w values are weights representing how important those inputs currently are. During training, the network keeps adjusting these weights when it makes mistakes.

Then comes an activation function such as ReLU (hero of Deep Learning):

def relu(x):
    return max(0, x)

So:

ReLU(-5) → 0
ReLU(-1) → 0
ReLU( 2) → 2
ReLU( 8) → 8

This looks almost too simple to be useful. But the magic is not in one ReLU or one neuron. It comes from connecting huge numbers of these mathematical operations together and stacking them into layers.

Very roughly:

Image
  ↓
Pixels
  ↓
[ Layer 1 ]  → simple patterns
  ↓
[ Layer 2 ]  → edges / shapes
  ↓
[ Layer 3 ]  → more complex features
  ↓
[ More layers... ]
  ↓
Dog: 97%
Cat: 2%
Other: 1%

This is obviously simplified, but it gives us the basic intuition.

One important distinction: our biological brain does not literally run ReLU functions. Artificial neural networks are mathematical systems loosely inspired by biological neurons. The analogy is useful, but they are not the same thing.

A tiny PyTorch example

In PyTorch, we could build an extremely simplified neural network like this:

import torch.nn as nn

model = nn.Sequential(
    nn.Linear(784, 128),
    nn.ReLU(),

    nn.Linear(128, 64),
    nn.ReLU(),

    nn.Linear(64, 10)
)
input 784 hidden + ReLU 128 hidden + ReLU 64 output 10
The network from the code above: 784 inputs flow through two hidden layers (128 and 64 neurons, each followed by ReLU) down to 10 outputs. Only a few neurons per layer are drawn; every line is a weight the network adjusts during training.

Don't worry too much about the syntax. Conceptually, we're saying:

784 inputs
    ↓
128 learned features
    ↓
ReLU
    ↓
64 learned features
    ↓
ReLU
    ↓
10 possible answers

Here is a fun bit of math. Even this toy network already has a surprising number of knobs to tune. Each layer has one weight for every input-output pair, plus one bias per neuron:

Layer 1: (784 × 128) + 128  = 100,480
Layer 2: (128 × 64)  + 64   =   8,256
Layer 3: (64 × 10)   + 10   =     650

Total: 109,386 parameters

The ReLU layers add nothing, since they have no learnable parameters. So this little example, small enough to fit on a slide, is already adjusting over a hundred thousand numbers during training. Keep that in mind for later, when we talk about models with billions of them.

During training, the model sees examples, makes predictions, compares them with the correct answers, calculates how wrong it was, and adjusts its weights a little. Repeat this millions or billions of times and the model gradually learns useful patterns.

Notice something important though: we never explicitly wrote those patterns into the code.

There isn't necessarily a programmer somewhere writing:

if ears == "floppy" and nose == "dog-like":
    confidence += 0.42

The network discovers useful internal representations by adjusting its weights during training.

That is where things start getting interesting.

From pixels to embeddings

Modern AI models go much further than simple image classifiers.

They learn representations of things.

Suppose we somehow represented words as coordinates:

dog    → [0.81, 0.72, 0.15, ...]
puppy  → [0.84, 0.75, 0.17, ...]
cat    → [0.76, 0.69, 0.20, ...]
car    → [0.12, 0.08, 0.91, ...]

These lists of numbers are called vectors, and learned vector representations are commonly called embeddings.

The individual numbers usually don't mean something simple like 0.81 = dog ears. What matters is the overall position and relationship between vectors. Concepts with related meanings can end up closer together in this mathematical space.

So instead of programmers manually defining:

dog.is_related_to("puppy")

the relationship can emerge from training.

Again, we didn't explicitly tell the model every relationship. It learned a representation from examples.

And then came Transformers

Large Language Models such as GPT add another major idea: attention.

Consider:

“The dog chased the ball because it was moving.”

What does it refer to?

Understanding language requires looking at relationships between words and their context. Transformers use an attention mechanism that, in simplified terms, allows the model to ask:

While processing this token,
which other tokens should I pay attention to?

A transformer does this across many layers and many attention heads. As information moves through those layers, the model builds increasingly useful internal representations of the context.

Scale this architecture, train it on enormous amounts of text, code, images and other data, and we get the foundation of today's large AI models.

Which brings us back to Polanyi

Here is the part I find most fascinating.

Suppose a large model correctly explains a difficult piece of code. We know mathematically what happened: tokens became vectors, those vectors passed through transformer layers, attention was calculated, matrices were multiplied, activation functions were applied, and eventually probabilities over the next token were produced.

We can inspect all of those calculations.

But ask:

“Where exactly inside those billions of parameters did the model learn the concept of a software bug?”

That becomes much harder to answer.

There isn't necessarily one neuron called:

neuron_847293 = "understands code bugs"

Knowledge can be distributed across many parameters, neurons, layers and interactions.

This is one reason AI interpretability is such an interesting research problem. We built the system. We know the mathematics. We can inspect every parameter. Yet understanding exactly why a large neural network arrived at a particular internal representation or answer can still be surprisingly difficult.

And that feels like an interesting echo of Polanyi's original observation.

I can look at a dog and immediately know it is a dog, but struggle to completely describe how I knew.

A neural network can classify the same dog, while we may struggle to translate millions of its internal computations into a simple human explanation of exactly how it knew.

The biological and artificial processes are very different, but the parallel is fascinating.

Maybe AI didn't solve Polanyi's paradox

This is probably my favorite way to think about it:

Machine learning didn't really solve Polanyi's paradox. It found a clever way around it.

Instead of forcing humans to explicitly describe everything we know, we started giving machines examples and letting them learn their own representations.

We couldn't write every rule for recognizing a dog, so we showed machines millions of images.

We couldn't write every rule of human language, so we trained models on enormous amounts of language.

We couldn't document every useful programming pattern, so models learned from huge amounts of code.

And today, with deep neural networks, embeddings, transformers, reinforcement learning and models containing billions or even trillions of parameters, machines can learn patterns that would be almost impossible for us to manually encode as rules.

That doesn't automatically make today's AI “superintelligent,” and scaling parameters alone isn't some magic formula for intelligence. But it does show how far we can get when we stop trying to tell machines every rule and instead build systems capable of learning patterns from experience and data.

More than sixty years later, Polanyi's simple observation feels strangely relevant to modern AI:

“We can know more than we can tell.”

And perhaps the more interesting question for the AI era is:

How much can machines learn from the things we were never able to fully tell them? 🐾