Four Options Were Making Me Feel Smarter Than I Was
Imagine the app shows me this word:
Obdurate
Then four options.
I look at them.
Three are obviously wrong.
One looks familiar.
I click it.
Correct.
The app could record that as success.
Maybe increase my mastery.
Maybe move the word slightly further away in the review queue.
Maybe tell me I know it.
There is just one problem.
What if I could not have produced the meaning myself?
What if I only recognised the answer because it was sitting in front of me?
What if I forget the word tomorrow?
What if I confuse it with another word that looks similar?
What if I understand the definition but cannot recognise the word inside an actual sentence?
Suddenly:
Correct answer
does not necessarily mean:
Known word.
That became the next problem with the vocabulary system.
And it forced me to change what vocabulary practice actually looked like.
The first vocabulary engine was already much better
This was not me discovering that flashcards are imperfect.
I had already moved past the original vocabulary-list version of the app.
The earlier system had become much more serious.
Words could have mastery.
The app could remember mistakes.
It could track lapses.
Review could happen over time.
Weak words could return.
Different kinds of evidence could influence what the app believed about my vocabulary.
That was already a huge improvement over:
Open card → read definition → mark learned.
But once that system existed, I started noticing another weakness.
Most of the evidence was still too easy to obtain.
Multiple-choice questions are useful.
They are also very good at making partial knowledge look complete.
Recognition is the easiest version of knowing
Suppose I see:
Laconic
and the options are:
- Talkative
- Brief
- Confused
- Angry
If I recognise brief, I probably know something about the word.
That evidence is useful.
But now remove the options.
Ask me:
What does laconic mean?
That is a different task.
There is nothing to recognise anymore.
The answer has to come from memory.
That difference matters.
Recognition gives the brain clues.
Recall does not.
And once I started thinking about vocabulary as evidence rather than checkboxes, the distinction became impossible to ignore.
A system that only tests recognition can become very confident about knowledge that disappears the moment the hints disappear.
I did not want that.
So I added typed recall
One of the biggest changes was active recall before reveal.
Instead of always giving myself the answer somewhere on the screen, the app could ask me to retrieve what I knew first.
See the word.
Try to remember it.
Type what I think it means.
Then reveal the actual answer.
That changes the experience immediately.
There is nowhere to hide behind good distractors.
Either something comes to mind or it does not.
And even when my answer is imperfect, the act of trying to retrieve it tells me something that tapping an option cannot.
Sometimes I would see a word and feel completely certain that I knew it.
Then the text box appeared.
Nothing.
That was useful.
Slightly painful.
But useful.
The word was familiar.
The meaning was not available.
Those are different states.
The app needed to understand that.
But recall is not the whole thing either
It would have been easy to swing too far in the other direction.
Multiple choice is imperfect.
Therefore everything should become typed recall.
That would create a different problem.
Vocabulary knowledge is not one-dimensional.
A student can recognise a word before they can recall it.
They can remember a rough meaning without understanding how the word behaves in context.
They can know two words separately and still confuse them when both appear together.
They can understand a root and use that structure to infer an unfamiliar word.
They can know a word perfectly and simply fail once because they were tired.
No single interaction can capture all of that.
So instead of trying to find the perfect vocabulary question, I started building multiple ways to interact with the same vocabulary.
The vocabulary system reached ten modes
At this point, the app had grown to 172 bundled words and expressions and 10 vocabulary learning modes.
That number is not important because ten is magically correct.
It matters because vocabulary practice no longer had to ask the same question in slightly different colours.
Different modes could test different kinds of knowledge.
Some could focus on recognition.
Some on recall.
Some on context.
Some on confusing alternatives.
Some on how words relate to other words.
Some on phrases rather than isolated dictionary entries.
The goal was not variety for entertainment.
It was variety because knowing a word has more than one failure mode.
Confusable words became their own problem
Some words are difficult because I do not know them.
Others are difficult because I know them badly enough to confidently choose the wrong one.
That second category is dangerous.
Words can look similar.
Sound similar.
Share roots.
Sit near each other in meaning.
Or simply become tangled together because I learned them at roughly the same time.
If the app treats every error as:
You forgot this word
then it misses something important.
Sometimes the real problem is:
You confused this word with that one.
So I started adding explicit confusable-word practice.
Instead of learning each item in isolation, the app could force me to distinguish between words that were likely to interfere with one another.
That gives the mistake much more meaning.
If I repeatedly choose one word when another is correct, the system is not merely observing failure.
It is observing a pattern.
And patterns are much more useful than isolated wrong answers.
Context exposed another kind of fake knowledge
Definitions are clean.
Language is not.
A word in a dictionary entry sits politely beside its meaning.
A word in a sentence has to survive grammar, tone and context.
So another part of the vocabulary system started testing whether I could recognise correct usage.
Not simply:
What does this word mean?
but:
Which sentence actually uses it correctly?
That sounds like a small change.
It is not.
Sometimes I know the definition and still misunderstand how the word should be used.
Sometimes multiple sentences look plausible.
Sometimes the surrounding context reveals that my understanding was too shallow.
This was exactly the kind of evidence the original cards could never produce.
The app knew I had seen the word.
Now it could begin testing whether my understanding survived contact with actual language.
Roots and families made vocabulary less isolated
Word Power Made Easy had already pushed me toward thinking about roots and related words.
That became increasingly useful once vocabulary stopped being a flat collection.
A word does not always need to be memorised as a completely independent object.
Sometimes understanding a root gives you a structure.
Related words become easier to connect.
Unfamiliar words become slightly less unfamiliar.
So roots and word families became part of the learning system too.
That does not mean every word can be solved through etymology.
It definitely does not mean recognising a root proves you know the word.
But it gives the app another teaching tool.
Instead of only saying:
Memorise this.
it can sometimes say:
Notice how this connects.
That is a much better experience for the right kind of word.
Idioms and phrasal verbs did not fit neatly into the old model
The vocabulary section was also becoming broader than single words.
Idioms.
Phrasal verbs.
Expressions.
These do not behave exactly like ordinary vocabulary entries.
Trying to force everything into:
word → definition
would have made the model simpler internally and worse educationally.
So the system needed to accept that different kinds of language may require different presentation and testing.
That sounds obvious now.
It was another example of the same lesson:
The data model should adapt to what I am trying to learn.
I should not distort what I am trying to learn because the original data model was convenient.
Then I needed a way to decide what to practise
Adding more modes created another problem.
If the app can test the same vocabulary in many ways, which mode should it use?
I did not want the student to open a menu every time and make ten decisions before learning one word.
That would be ridiculous.
The whole point of the project was to reduce friction.
So the vocabulary engine needed to start choosing more intelligently.
Not randomly.
And not with one universal sequence where every student receives the exact same activity in the exact same order.
The current system can use the evidence it already has to decide what kind of interaction makes sense next.
A new word may need one kind of exposure.
A weak word may need another.
A recently forgotten word should probably not be treated like something completely new.
A word that keeps getting confused with another deserves different practice from one that simply has not appeared for a while.
I deliberately do not want one number pretending to explain all of that.
The important part is that review now has a reason.
"Because the algorithm said so" is not good enough
This became surprisingly important.
If a word returns, I want the system to have some understandable reason.
Due for review.
Recently forgotten.
Weak recall.
Repeated confusion.
Needs more evidence.
Whatever the reason is, there should be one.
That makes the vocabulary queue easier to trust.
Otherwise adaptive software can start feeling arbitrary.
A card appears.
Why?
Who knows.
The number moved.
Why?
Algorithm.
I do not want that.
The system can be complicated underneath.
The learning experience should still make sense from the outside.
I also built a placement diagnostic
There was another obvious problem.
What happens when someone already knows a lot of the vocabulary?
Treating every word as brand new would waste time.
But assuming too much would be equally bad.
So I added a small placement diagnostic.
Currently, it uses 12 words to gather an initial signal.
Not to declare:
This is your permanent vocabulary level.
That would be absurd.
Twelve words cannot do that.
The purpose is much smaller.
Give the engine somewhere better to start than zero information.
The diagnostic can provide an initial indication of what appears familiar, what does not, and where the system may need more evidence.
Then normal use can gradually replace that first estimate with actual behaviour.
That is a pattern I like more and more:
Start with a useful guess.
Then earn confidence through evidence.
The system now tracks more than one kind of vocabulary signal
This also meant the app could no longer think about vocabulary as one percentage.
Imagine seeing:
Vocabulary: 72%
What does that actually tell me?
Maybe I am excellent at recognising definitions and terrible at recall.
Maybe I remember individual words but struggle with context.
Maybe I repeatedly confuse similar-looking terms.
Maybe I know common words but collapse on phrases.
One percentage hides all of that.
So vocabulary started feeding multiple kinds of signals into the broader learning system.
The goal is not to create a dashboard with twenty impressive numbers.
It is to avoid making one number carry meaning it does not deserve.
If different weaknesses require different repair, the system needs enough structure to tell them apart.
Daily practice also became more deliberate
At this point, vocabulary had enough moving pieces that opening the section and choosing manually every time would have been annoying.
New words.
Due words.
Weak words.
Recent mistakes.
Items that had not appeared for a while.
Harder material.
Different modes.
So I added a mixed daily mission.
One button.
Start.
The session can combine different kinds of useful work instead of asking me to construct the perfect study queue myself.
That idea matters to me because one of the easiest ways to avoid studying is to turn preparation into planning.
Choose the perfect mode.
Choose the perfect list.
Decide what deserves review.
Reorganise everything.
Twenty minutes later, congratulations.
No vocabulary learned.
The app already has information about what happened previously.
It should use that information to reduce the decisions required before I begin.
Recently forgotten words deserved special treatment
One of the more useful changes was becoming more sensitive to lapses.
There is a difference between:
I have never learned this
and:
I knew this and just forgot it.
The second case contains history.
The system already has evidence that the word was once stronger.
Then something changed.
That is valuable.
A recently forgotten word may deserve attention sooner than a completely untouched item.
Repeated lapses may suggest that the current learning pattern is not holding.
Again, I do not want to publish or obsess over one magical formula.
The important design principle is simpler:
Forgetting is evidence too.
Not a failure state to erase.
Not something that disappears the moment I answer correctly once.
Something the system can learn from.
This is where spaced repetition became more meaningful
It would be easy to say:
I added spaced repetition.
Done.
But spacing alone does not solve the problem.
You still need to decide what evidence changes the schedule.
What counts as success?
What happens after a lapse?
Does one correct recognition answer move the word far away?
Should strong recall count differently?
What if the student keeps confusing the same pair?
How much confidence should the system have after only a few interactions?
Those questions are much more interesting than simply attaching dates to cards.
So the review system became more retention-aware.
Not because I wanted to invent some mystical memory algorithm.
Because the app needed to stop behaving as if all correct answers meant the same thing.
Ten modes can still become ten gimmicks
This was something I had to keep reminding myself.
Adding ten learning modes sounds impressive.
It can also become the vocabulary version of old EBO.
More features.
More buttons.
More apparent sophistication.
No actual improvement.
So the number of modes is not the achievement.
The test is whether each mode reveals or teaches something the others do not.
If two modes produce basically the same evidence, I probably do not need both.
If a mode is entertaining but does not improve learning, that is not enough.
If an interaction adds friction without helping the system understand anything new, it should be questioned.
EBO already taught me what happens when I confuse quantity with product quality.
I do not want to relearn that lesson inside another project.
There are still things the vocabulary system cannot do
This matters.
The vocabulary engine is much better than the version I started with.
It is not complete.
One major gap is productive language.
Recognising a correctly used sentence is useful.
Actually producing one yourself is harder.
The app is not yet at the point where I would claim that it can reliably evaluate open-ended vocabulary usage.
That kind of feature sounds very impressive.
It also creates much more room for bad feedback.
So I would rather call the gap what it is than put an "AI writing evaluator" badge on the interface and pretend the problem is solved.
There are other limitations too.
The vocabulary library is still growing.
The content still needs review.
Different modes will need more real use before I know which ones are genuinely valuable.
The diagnostic is an initial signal, not a psychological assessment.
Adaptive decisions can become better as the app gathers more evidence.
This is still a system under construction.
The strange part is that the app now knows less confidently
That sounds backwards.
The old system could say:
You got it right.
Very clean.
Very certain.
The newer system asks:
Did you recognise it?
Could you recall it?
Did you understand it in context?
Have you confused it before?
When did you last see it?
Did you recently forget it?
How much evidence do we actually have?
That produces messier answers.
But I trust them more.
The goal was never to make the app sound certain.
The goal was to make its confidence deserve to exist.
Sometimes:
Not enough evidence yet
is a much smarter conclusion than:
Mastered.
Four options were making me feel smarter than I was
I still use multiple-choice questions.
They are useful.
They resemble parts of the exams I am preparing for.
They can test recognition efficiently.
They can expose good distractors and common confusions.
The mistake was never using MCQs.
The mistake was allowing one kind of success to represent the entire concept of knowing.
Vocabulary is messier than that.
So the app became messier too.
Typed recall.
Context.
Confusables.
Roots.
Families.
Idioms.
Phrasal verbs.
Multiple evidence signals.
Adaptive review.
A placement diagnostic.
A mixed daily mission.
More ways for the system to discover that I do not know something as well as I thought I did.
Which is slightly rude.
But useful.
The first vocabulary version helped me see words.
The next version helped me review them.
The current one is starting to ask a much harder question:
What evidence would actually convince me that this word is staying in my head?
Four options were not enough.
Apparently building ten different ways to prove myself wrong was the next logical step.