I Stopped Building a Vocabulary List and Built a Vocabulary Engine
For a while, I thought the vocabulary part of this app was already pretty good.
That was probably because I was comparing it to where it started.
Originally, vocabulary was barely a system at all.
There were words.
There were definitions.
Eventually there were 154 vocabulary cards.
I could browse them.
Learn them.
Review them.
Vocabulary activity could count toward my streak.
Compared with the first version of the project, that felt like a huge improvement.
And it was.
But after using it for a while, I started noticing a problem.
The app knew that I had seen words.
It did not necessarily know whether I actually knew them.
Those are very different things.
Seeing a word is not learning a word
Suppose the app shows me:
Obdurate
I read the definition.
Maybe I recognise it.
Maybe I press something that tells the system I understood it.
Great.
What has actually been proven?
Almost nothing.
Could I recognise the word tomorrow?
Could I distinguish it from a similar word?
Could I choose its meaning inside an MCQ?
Could I understand it inside a sentence?
Could I identify the correct synonym?
Could I avoid a convincing wrong option?
Could I still remember it a week later?
A vocabulary app can create the illusion of progress extremely easily.
Show enough cards.
Increase a counter.
Display a streak.
Tell the student they learned 40 words.
The numbers go up.
Everybody feels productive.
And then the exam asks the word in a slightly unfamiliar way and the entire illusion collapses.
I did not want my app doing that.
I needed a better definition of "known"
That became the real problem.
What does it mean for the system to say:
Mikhil knows this word.
It cannot simply mean:
"He opened the card once."
It probably should not even mean:
"He answered one question correctly."
One correct answer can happen because the options were obvious.
Or because I guessed.
Or because I had seen the word thirty seconds earlier.
Actual vocabulary strength has more dimensions.
Recognition.
Recall.
Context.
Confusability.
Retention over time.
The ability to recover after forgetting something.
Once I started thinking about vocabulary this way, the existing system suddenly looked much more primitive.
The cards were useful.
But cards were only an interface.
I needed an engine underneath them.
So vocabulary became its own milestone
Until this point, I had mostly been developing the entire verbal app together.
Practice.
Review.
Progress.
Offline behaviour.
Vocabulary.
Everything evolved alongside everything else.
But the project had become complicated enough that continuing like that was starting to feel messy.
So I began treating major parts of the app as separate milestones.
Vocabulary became one of them.
The goal was not:
Add more vocabulary features.
It was:
Build the strongest vocabulary learning system I reasonably can for this app.
I started calling it the Peak Vocabulary Engine.
The name is slightly dramatic.
But honestly, dramatic milestone names make building things more fun.
The app should care about how I know a word
One major change was moving away from the idea that every vocabulary interaction is equivalent.
A student can be tested in different ways.
Direct meaning recognition is one.
Context is another.
Synonyms and antonyms test something slightly different.
Confusable words create another problem entirely.
If I repeatedly mix up two similar-looking or similar-sounding words, showing me the same definition card again probably is not enough.
The system needs to understand the failure.
That changed how I thought about vocabulary questions.
A question was no longer just:
Did he get this right?
It could also tell the system something about what kind of knowledge failed.
That is much more useful.
Wrong answers became evidence
This is a pattern that keeps repeating throughout the project.
At first, wrong answers were things I wanted to save for later.
Then they became something more valuable.
Evidence.
Suppose I get a word wrong once.
That might not mean much.
Everybody makes mistakes.
But if I fail the same word repeatedly, something is clearly happening.
If I understand its direct definition but keep missing it in context, that tells us something else.
If I consistently confuse it with another word, that is another signal.
A useful learning engine should not react to all of those situations identically.
So vocabulary review started becoming less about maintaining a giant queue and more about interpreting what happened.
Review should have a reason
This is probably the biggest philosophical change.
A bad review system says:
You learned this before. Here it is again.
A better system asks:
Why should this word appear now?
Maybe it has been long enough that retention should be checked.
Maybe I failed it recently.
Maybe I have not demonstrated enough understanding yet.
Maybe I keep confusing it with something else.
Maybe I repaired it once but then lapsed again.
Maybe it is already strong and does not deserve more of my time right now.
That final one matters too.
Review systems can waste time just as easily as they can save it.
If I already know something extremely well, repeatedly showing it to me because it happens to be in the database is not personalised learning.
It is noise.
Mastery needed to be earned
That meant the app needed a stronger concept of mastery.
Not:
Card opened = mastered.
Not:
One correct answer = mastered.
Mastery should represent accumulated evidence.
Enough successful interactions.
Enough variety in how the word was tested.
Enough time to show that I did not simply remember it for thirty seconds.
And even then, mastery should not necessarily be permanent.
People forget.
I definitely forget.
A word I knew confidently two weeks ago can suddenly look like it was invented five seconds ago.
So the engine also needed to understand lapses.
That sounds negative, but I actually like the idea.
If I forget something, the app should not pretend my old mastery still means I am fine.
It should respond.
Bring the word back.
Repair the weakness.
Then let it become strong again.
This made progress more honest
One thing I increasingly dislike in educational software is fake certainty.
A progress bar reaches 100 percent.
A unit turns gold.
A word gets a green tick.
And the system acts like knowledge has permanently entered your brain.
Real learning is much messier.
You learn something.
You strengthen it.
You forget part of it.
You recover it.
Some things become effortless.
Some remain fragile for months.
I wanted the vocabulary engine to reflect that reality better.
That means the progress numbers might become less flattering.
But they become more useful.
I would rather have the app tell me:
You still do not reliably know this.
than congratulate me because I looked at it once.
Daily vocabulary also became more deliberate
The original vocabulary system already contributed to daily activity.
Reviewing or learning enough words could help maintain the streak.
But once the engine became smarter, I started thinking about the daily vocabulary experience differently too.
I do not necessarily want to open the app and stare at the entire word collection.
I want a useful amount of work for today.
Words worth learning.
Words worth reviewing.
Weak vocabulary worth repairing.
A manageable mission rather than an endless library.
The giant library can still exist.
But the app should help decide what deserves attention now.
That distinction is becoming important across the entire project.
I am less interested in giving the student access to everything.
I am more interested in helping them choose what matters next.
Confusable words deserved special treatment
Some vocabulary failures are particularly annoying.
You know both words.
At least you think you do.
Then the exam places them together and suddenly your brain merges them into one object.
Those pairs are much more dangerous than a completely unfamiliar word.
If I have never seen something before, I know I do not know it.
Confusable words create false confidence.
So they deserve more direct confrontation.
Show them together.
Force the distinction.
Test the actual difference.
Do not let me escape by memorising two isolated definitions separately.
That became another place where the engine could be much more useful than a static word list.
The goal was not to add AI everywhere
Whenever a system starts adapting to the user, it becomes tempting to call everything AI.
I am trying not to do that.
I do not care whether the vocabulary engine sounds technologically impressive.
I care whether it makes better decisions.
Does it know what I should review?
Does it stop wasting time on strong words?
Does it bring back fragile ones?
Does it recognise repeated failure?
Does it treat confusable vocabulary differently?
Does it give me enough variety to prove I actually understand something?
If the answer is yes, the system is useful.
The label is secondary.
The content problem is still enormous
Building a smarter engine created an ironic problem.
The better the system becomes, the more obvious the limitations of the content become.
An adaptive vocabulary engine with fifty words would be pointless.
Even 154 cards are nowhere near the final ambition.
The engine needs enough high-quality vocabulary to actually have something useful to schedule.
And the vocabulary cannot just be random.
This is still an IPMAT and IIM UG preparation product.
Relevance matters.
Difficulty matters.
The kinds of questions attached to each word matter.
Quality matters.
So finishing the engine did not mean vocabulary was finished.
It meant the infrastructure was finally becoming capable of handling the vocabulary system I actually want to build.
The content can now grow into it.
This is why I stopped adding thousands of questions immediately
Around this stage, I also had to make a decision about the larger question bank.
I could stop everything and spend a huge amount of time producing content.
Or I could finish the major engines first, make sure the product architecture was strong, and then scale the question bank systematically.
I chose the second option.
The plan now is to build questions in manageable batches.
Roughly 250 at a time.
Review them.
Integrate them.
Then continue.
The initial serious target is around 2,500 questions.
That is already enough content to make the app substantially more useful without pretending the database is somehow finished forever.
And unlike the original 32-question prototype, the system those questions enter is now far more capable.
The difference is becoming obvious
The first vocabulary version answered:
Which words are in the app?
The version I am building now tries to answer:
Which words should I work on?
Why am I seeing this one again?
Do I actually know it?
What kind of mistake am I making?
When should it return?
What should happen if I forget it again?
That is a completely different product problem.
And it made me realise something.
When I started this project, I thought the hard part would be collecting enough vocabulary.
Now I think collecting words is probably one of the easier parts.
The harder problem is deciding what the app should do with them.
The vocabulary milestone is done
At some point during this process, I looked at the system and realised I had crossed an important line.
There were still more words to add.
More questions to create.
More visual polish I could do.
There always will be.
But the underlying vocabulary milestone itself finally felt complete enough that I could stop adding random improvements and move forward.
That mattered because there was another major promise from the previous version of the app that I still had not properly delivered.
I had talked about Learn mode.
Paced practice.
Strict exam conditions.
Real timing.
Actual exam structures.
Until now, most of that existed as product direction.
The vocabulary engine was finally strong enough for me to move on.
So the next milestone became obvious.
It was time to stop talking about an exam mode.
And build the actual exam engine.