Why I built Verba

I’ve tried every tool and every strategy that people recommend for language immersion. I’ve spent countless hours consuming content where I just couldn’t make sense of what was being said. And the advice that was given to me was trust the process, tolerate the ambiguity, and eventually I just got burned out.

From many conversations with people in the community I realised this is quietly experienced by a lot of people. I believe there is a survivorship bias in the immersion learning community: those who are successful are held up as examples of the effectiveness of the method, while those who fail to reach fluency are led to believe their failure was a personal failing rather than a problem with the method.

Survivorship bias: a stick figure on the tip of a pyramid says 'If I can do it, anyone can!' The tip is the 0.1% who succeeded. The rest is the 99.9% who failed.

Learning a language to a high level will always be something that takes time to achieve, at least until we invent that technology from The Matrix where you can upload knowledge directly to your brain. Until then we are stuck with the reality that learning a language involves encoding millions of connections between neurons.

Neo from The Matrix with the subtitle 'I know kung fu', where kung fu is crossed out and replaced with Japanese

But does it have to be as slow and painful as it is right now? I’ve spent a long time thinking about this question.

Over the last few years, a large chunk of my time has been devoted to answering it: poring through linguistics research, talking to professors from Princeton and MIT about how language acquisition works, testing every product on the market.

I learned that linguists actually have a pretty clear picture of the exact conditions under which language acquisition occurs. And unless you’re optimising for those conditions you can end up spending a lot of time not learning anything.

So what are those conditions? Let me walk you through a set of principles about how language acquisition occurs according to research from the best linguists alive right now.

How language acquisition works (according to science)

1. We don’t learn words as definitions. We learn them as cues, based on the patterns we’ve seen them in.

It’s intuitive to think of words having definitions. If I ask you what the word “run” means, you’ll likely think of the most common usage, which is to move your legs fast to move fast. But there are actually endless ways to use the word run:

  • Run away (flee)
  • Run after someone (chase)
  • Run off with something (leave while taking what is not yours)
  • Run for office (stand as a candidate in an election)
  • In the running (still having a chance of winning)
  • Score a run (a point scored in cricket or baseball)
  • Run a business (manage and be in charge of)
  • Run a program (start a process and keep it going)
  • Up and running (working or functioning)
  • Run on diesel (use as a power source)
  • The tap is running (liquid flowing)
  • The colours ran (dye spreading out of its area in the wash)
  • A run in her tights (a line of unravelled stitches)
  • Run a fever (have a raised temperature that persists)
  • The road runs along the coast (extend in a continuous line)
  • A ski run (a marked course to travel down)
  • … it goes on and on.

All of these are quite distinct concepts that are only loosely connected through metaphor, and you instantly understand their meaning because at some point you learned that usage pattern.

Which means you didn’t learn the word “run” once. You actually learned hundreds of versions of the word that mean all kinds of things.

Which is why vocabulary lists aren’t effective. Words don’t exist in our heads like entries in a database. Instead our brains are aware of many, many usage patterns that we learn from exposure.

2. To learn a new word, you have to see it not just multiple times, but in varied and informative contexts

Our brains learn through prediction and error correction. As you’re reading or listening, you pre-empt the meaning of a word based on similar contexts you’ve seen it in. If you were right, the pattern gets reinforced. If you were wrong, your brain has something to correct and starts detecting this new pattern next time it encounters it.

Without variation this cannot happen. It is exposure to varied usage patterns that allows the brain to triangulate meaning through making predictions.

The other part that’s important is that the context is informative. Essentially, if you cannot easily infer meaning from context (prediction) then the learning process cannot occur.

So imagine you’re learning English, you’re unfamiliar with the word “run” and you see the phrase, “She’s been running since March.” Depending on the surrounding context, it can be really hard to infer what “running” actually means here: training for a marathon, carrying out an election campaign, or keeping a business open. Whereas the sentence, “She’s been running the cafe since March,” gives you a lot more to work with and makes it a lot easier for your brain to detect meaning from the usage pattern.

A marathon runner wearing a 'Vote for me for mayor' sash while carrying a cash register and a calculator

Which is why you can’t truly learn a word using flashcards. A card is a static context, so the tenth review is the same data point as the first. You put the word “run” on a flashcard with the sentence “He runs every morning” on the back of the card. You review it 100 times and you’ve only learned one thing. You only learned one facet of how that word is used. You haven’t really learned the word.

3. The brain doesn’t learn without attention

Encoding new information in the brain cannot happen unless that information was actively processed.

Which means if you are zoning out because you don’t understand, or you’re bored from grinding flashcards, you likely aren’t learning much at all.

Meme of a purple lizard staring blankly, captioned 'Are u listening' and 'Me zoned out:'

4. Progress slows dramatically once you get past beginner words

There’s an interesting phenomenon in all languages, called Zipfian distribution. If you list all the words in a language by frequency, you’ll find that the second most common word turns up about half as often as the first, the third about a third as often, and so on all the way down.

How often you encounter words from different frequency bands:

  • A word in the top 1,000 most frequent words appears around 7 times per hour in normal speech.
  • A word between the 3,000th and 4,000th most frequent words appears only once per 8 hours of ongoing speech.

Sources: Nation (2006), Webb (2007).

Zipf curve: word frequency drops steeply after the most frequent word, then flattens into a long tail

Which means that if you’re doing 1 hour of immersion a day, and given it takes on average 10 exposures to learn a word, a word in the top 1,000 you can learn in 90 minutes of listening. However, a word ranked between 3,001 and 4,000 would take about 3 months for you to get those 10 exposures, and that’s just for one word!

What is the ideal immersion environment?

Most people are optimising for two things:

  1. Number of hours spent immersing.
  2. Anki reps consistency.

Given what I’ve outlined above, it should be clear that optimising for hours and reps is actually a brute force approach to language acquisition. The hope is that you will eventually get enough exposures in enough contexts that you build up a native-like model of the language in your brain (and don’t get burned out trying to reach fluency).

However, due to Zipfian distribution, once you get past beginner level it becomes increasingly difficult not just to get enough exposure to the words you’re learning, but to get it in contexts that you can effectively learn from.

For a long time I’ve dreamed about the ideal immersion environment, which is one where you are always exposed to words in ideal learning contexts at the precise time you need to see them (which I call effective exposure).

What attributes do ideal learning contexts have?

  • Varied (so the brain can learn from prediction)
  • Informative (inferable)
  • Engaging (so you’re paying enough attention)
  • Shown to you at the right time (so you don’t forget)

And this is what I set out to build: a content platform that maximises effective exposure.

So what is Verba?

Verba is what I’m calling an Optimised-Immersion Platform. It collapses the boundary between SRS flashcards (like Anki) and comprehensible input immersion.

The most effective approach for language learning right now is to combine immersion (in natural input) and Anki flashcards.

As I talked about before, once you get past beginner level it becomes increasingly difficult to learn new words, because the words you need to learn don’t show up often enough in the input.

SRS flashcards are meant to solve that problem by showing you words you are learning when you need to see them, so you don’t forget them (because you don’t see them often enough in the input). The problem is that you can’t truly learn a word with flashcards. It can only train you to get better at noticing it.

Which means the word exposure problem isn’t really solved.

When you think about what Anki is doing, it’s just showing you a word when you need to see it.

What if during your immersion, the words that you need to see at that time always show up in the content? Anki would become unnecessary. Verba does exactly this.

Using a context algorithm, Verba recommends content to you based on two metrics:

  1. whether it contains enough words that are due for review according to SRS (Verba uses the same SRS algorithm as Anki)
  2. whether it contains enough words that you have previously interacted with

I call this Optimised Immersion.

At the top of the recommendations page is a collection of videos that the context algorithm has found to contain the most due and seen words:

The Verba recommendations page. A row called 'Review your due words' shows videos, each with a count of your due reviews and previously seen words.

The metrics for each of these videos are shown below:

Close-up of the metrics under a recommended video: contains 78 of your due reviews and 18 previously seen words

Words that are due for review are highlighted in the transcript:

A Japanese video transcript in Verba. Words that are due for review are outlined in green.

Clicking a due word will show the flashcard for that word:

The flashcard for こんにちは, opened from the transcript. It shows audio, the translation, the dictionary form, and Forgot and Remembered buttons.

So no more grinding through endless Anki reps. With Verba, all you have to do is immerse, because Verba is designed to find the content that contains the words you need to see.

The added benefit to this is, of course, that every time you review one of the words you’re learning, you’re seeing it in a fresh context.

What Verba does not do (yet)

In future, the context algorithm will recommend content not just based on due words and seen words, but also on whether it fits your interests and whether you are likely to find it interesting. But right now, it picks from a growing library of YouTube videos: a mixture of comprehensible input content and random videos on different topics, news, fashion, vlogs, etc.

Another metric a future version of the algorithm will index for is informative contexts, which I discussed before: contexts where it’s easy to understand and infer how a word is really being used and what it really means.

Right now, of the four components of effective exposure, Verba delivers on two of them (variation and timing). Informative contexts and recommendations based on your personal interests are coming later.

How to use Verba

Right now Verba works with Korean and Japanese videos from TikTok or YouTube.

If you are on the Everyday plan, you’ll be able to import videos from the recommendations.

If you are on the Max plan you can also import your own videos:

The import bar in Verba: a 'Paste link' field for a YouTube or TikTok URL, an Import button, and the library list below

Verba will take a bit of time working through the video, and do the following things:

  1. Transcribe it from scratch (far more accurate than YouTube auto-subs)
  2. Parse and process all the sentences, chunks and words
  3. Add it to your library so you can work through the video any time you like

So import a couple of videos, and while you’re waiting for them to process, work through some videos from the Recommended page.