Back to all writing

My App Started Telling Me What I Was Bad At

IPMAT Series•Part 9•11 min read•By Mikhil

My App Started Telling Me What I Was Bad At

By this point, the app knew a lot about me.

Probably more than I realised when I started building it.

It knew which questions I had attempted.

Which ones I got wrong.

Which words kept coming back.

How long I took.

What happened inside mocks.

Which topics appeared repeatedly.

How consistent I had been.

Which mistakes I repaired.

Which ones returned.

That sounds like a lot of useful information.

And it is.

But there was a problem.

Most study apps stop at collecting data.

They show you charts.

Percentages.

Streaks.

Accuracy numbers.

Time spent.

Maybe a nice heatmap.

Then they leave you to figure out what any of it means.

I did not want to build another dashboard that says:

Grammar accuracy: 61 percent

and then expects me to become my own analyst.

If the app already has the evidence, it should help interpret it.

That became the next milestone.

Personal Intelligence.

Analytics are not intelligence

This distinction became important very quickly.

Suppose the app tells me:

You answered 68 percent of vocabulary questions correctly.

That is analytics.

Useful, but limited.

Now suppose it tells me:

Your direct vocabulary recognition is strong, but you are repeatedly losing accuracy when words appear in context.

That is much closer to intelligence.

The first system reports what happened.

The second tries to interpret the pattern.

That is what I wanted.

Not some magical AI tutor pretending to understand my entire brain.

Just a system that takes the evidence it already has and turns it into something actionable.

The app needed to know where I was weak

One of the first obvious additions was a skill heatmap.

Not just one giant verbal score.

That number hides too much.

If I am excellent at reading comprehension and terrible at grammar, averaging them together gives me a score that describes neither problem properly.

So the app needed to break performance down.

Vocabulary.

Grammar.

Reading.

Paragraph skills.

And eventually smaller subskills inside them.

The idea is simple:

Show me where the evidence is strong.

Show me where it is weak.

Show me where there is not enough evidence yet.

That last part became surprisingly important.

Two wrong answers should not diagnose me

Imagine I answer two questions from a topic.

I get both wrong.

The easiest possible system immediately says:

Weakness detected.

That looks intelligent.

It is also ridiculous.

Two questions are not enough evidence to confidently diagnose anything.

Maybe I was distracted.

Maybe both questions were unusually difficult.

Maybe I guessed.

Maybe the sample simply happened to go badly.

So I added evidence thresholds.

The system should not confidently tell me:

You are weak at this

until it has enough information to justify the claim.

Before that, the honest answer is closer to:

Not enough evidence yet.

I like that much more.

A system that admits uncertainty is more useful than one that confidently invents precision.

Weakness has more than one shape

Even when the app has enough evidence, poor performance can mean different things.

Suppose my accuracy is low.

Why?

Maybe I genuinely do not understand the topic.

Maybe I understand it but answer too quickly.

Maybe I spend too long and still get it wrong.

Maybe I repeatedly fall for the same distractor pattern.

Maybe I learned something previously and then forgot it.

Maybe my performance is fine in normal practice but collapses during a timed mock.

Those are not the same problem.

So simply labelling something:

Weak

is not enough.

I wanted the app to start classifying what kind of problem might be happening.

Mistakes became categories

Wrong answers had already evolved several times throughout this project.

First they were just mistakes.

Then they entered review.

Then they became evidence for vocabulary scheduling.

Now I wanted to classify them more meaningfully.

The exact system can keep evolving, but the important idea is that mistakes should have structure.

A careless error is different from a knowledge gap.

A timing problem is different from misunderstanding the concept.

Repeated confusion is different from seeing something for the first time.

If the app can separate those patterns, then the solution can become different too.

That is when analytics starts becoming useful.

Speed and accuracy needed to be seen together

The exam engine made timing data much more important.

Looking at speed alone is dangerous.

Fast is not automatically good.

Slow is not automatically bad.

Suppose I answer vocabulary questions extremely quickly but my accuracy is terrible.

The problem is probably not:

Become even faster.

Likewise, if I am highly accurate but take far too long, then accuracy alone hides the issue.

So the intelligence layer started looking at timing alongside correctness.

That creates more useful patterns.

Fast and accurate.

Fast and inaccurate.

Slow and accurate.

Slow and inaccurate.

Those combinations tell very different stories.

And they can lead to very different recommendations.

Mocks became evidence instead of isolated scores

This was one of the most important consequences of building the exam engine first.

Previously, a mock could have ended with:

Score.

Accuracy.

Time.

Done.

Now the attempt could feed the broader learner model.

Maybe a topic looks fine during casual practice but performs badly under strict timing.

That is interesting.

Maybe accuracy stays high but response time becomes unstable.

That is interesting too.

Maybe I leave too many questions unattempted in one category.

Or spend too much of the exam on a small number of difficult questions.

A mock should not become a dead historical record.

It should change what the app thinks about my preparation.

Retention became part of the picture

The vocabulary engine had already introduced a stronger idea of forgetting.

That concept also belongs in the wider intelligence system.

I care about whether I can answer something today.

But I also care about whether I can still answer it later.

A topic that looks strong immediately after practice may not actually be stable.

A word I repeatedly relearn is different from a word that stays mastered.

So the app began tracking retention-related signals too.

Not because I need another graph.

Because forgetting is evidence.

If the app sees that something repeatedly decays, it should react differently from something that remains stable.

I wanted weekly summaries that actually say something

A weekly report sounds like a very normal feature.

You studied for this long.

You answered this many questions.

Your streak is this number.

Fine.

But again, those are mostly counts.

I wanted the weekly summary to answer better questions.

What improved?

What became worse?

Where is there now enough evidence to identify a weakness?

Which mistake keeps repeating?

What seems stable?

What needs repair next?

What should I probably focus on this week?

That is much more useful than:

You answered 86 questions. Great job!

I already know I answered them.

Tell me what happened because of it.

Then came recommendations

Once the app can interpret performance, recommendations become possible.

This is also where things can become stupid very quickly.

A recommendation system that simply says:

Practice grammar

because grammar has the lowest percentage is not especially intelligent.

The recommendation should depend on evidence.

Maybe the topic has enough attempts.

Maybe the weakness is recent.

Maybe the problem is persistent.

Maybe the student has already been repairing it.

Maybe another area is currently more urgent.

I wanted recommendations to feel earned by the data rather than randomly generated from a list of motivational phrases.

The app does not need to pretend to be a teacher.

It just needs to make a better next-step suggestion than:

Do more questions.

Repair sets were the obvious next step

Eventually, I realised that telling me about a weakness and then making me manually find relevant questions was unnecessary friction.

If the system says:

This area needs work

then it should be able to help me work on it.

That led to repair sets.

A repair set is essentially targeted practice assembled around something the app thinks deserves attention.

Repeated errors.

Weak topics.

Fragile skills.

Things that need another attempt.

Instead of:

Dashboard → notice problem → remember problem → navigate elsewhere → configure session → find relevant questions

the app can shorten that whole chain.

You are struggling here. Repair it now.

That is much closer to the product I originally imagined.

This is where the app started choosing for me

Not completely.

I still want control.

I should be able to choose a topic.

Choose a session length.

Take a mock.

Learn vocabulary.

Browse content.

But there is a huge difference between allowing choice and forcing the student to make every decision manually.

When I open the app tired, I do not necessarily want to diagnose my own preparation.

Sometimes I want the system to say:

Do this next.

And have a good reason for saying it.

That was the point where the product began feeling genuinely different from the first prototype.

The original app served questions.

This version could begin deciding which questions were worth serving.

I had to resist making the dashboard ridiculous

There is a dangerous moment when you have a lot of data.

You want to show all of it.

Heatmap.

Accuracy chart.

Timing graph.

Retention chart.

Topic breakdown.

Difficulty breakdown.

Weekly comparison.

Streak.

XP.

Mock history.

Mistake distribution.

Recommendations.

Everything.

At once.

Because if the data exists, surely the user should see it.

I do not think that is true.

A dashboard can become a spreadsheet wearing nice colours.

The goal is not to prove how much information the app stores.

The goal is to help me understand what matters.

So the intelligence layer also forced me to think about hierarchy.

What deserves attention now?

What can stay in deeper analytics?

What is actionable?

What is merely interesting?

That problem is not completely solved.

But at least I know it exists.

The app also needed to avoid pretending certainty

This became almost a design principle.

If evidence is weak, say so.

If the system has only seen a few attempts, do not diagnose aggressively.

If two signals conflict, do not hide the uncertainty.

If there is no meaningful recommendation yet, do not manufacture one just to fill a card.

I think educational software can become dangerous when it sounds more certain than the evidence supports.

Especially once students begin trusting its recommendations.

So I would rather have a blank state than fake intelligence.

This milestone made earlier systems more valuable

One thing I really like about the way the app has evolved is that new systems make old systems more useful.

The question bank was originally just content.

Then the practice engine used it.

The exam engine used it differently.

Now the intelligence layer learns from both.

Vocabulary reviews were originally just revision.

Now they also produce retention evidence.

Timing was originally just a number.

Now it can help distinguish different performance problems.

Wrong answers were originally just things to retry.

Now repeated mistakes can influence recommendations and repair sets.

The systems are beginning to connect.

That is where the app starts feeling like one product rather than a collection of features.

It finally started doing what I wrote about earlier

A few articles ago, I wrote that I did not want random practice forever.

I wanted the app to eventually understand enough about my preparation to decide what would be useful next.

At the time, that was still mostly a direction.

Now the foundation for that actually exists.

Heatmaps.

Error patterns.

Timing.

Retention.

Evidence thresholds.

Weekly summaries.

Recommendations.

Repair sets.

None of those individually are revolutionary.

The interesting part is how they work together.

The app sees what I do.

Stores the evidence.

Looks for patterns.

Avoids making claims when the sample is too small.

Then uses stronger patterns to influence what it shows me next.

That is the first version of Personal Intelligence I actually wanted.

And then I had to make sure I could trust all of it

There was one uncomfortable consequence.

The smarter the app became, the more damaging unreliable data would be.

If one attempt disappears, the dashboard changes.

If offline progress gets overwritten, the learner model becomes wrong.

If a sync race duplicates an action, analytics become distorted.

If an update breaks cached content, practice can fail.

If I change devices, my learning history needs to survive.

Personal Intelligence only works if the information underneath it is trustworthy.

That meant the next milestone was not glamorous.

No flashy new learning engine.

No exciting exam mode.

No clever recommendation system.

Instead, I had to work on the kind of features nobody cares about until they fail.

Offline queues.

Sync safety.

Content packs.

Backups.

Updates.

Account controls.

Deletion.

Basically:

Can I trust this app with the history it is now using to understand me?

That became the next thing I had to solve.