Back to all writing

This Started With 32 Questions. Here’s What It Became

IPMAT Series•Part 17•14 min read•By Mikhil

This Started With 32 Questions. Here’s What It Became

The first version had 32 questions.

I keep coming back to that number because it makes the current project slightly ridiculous.

Thirty-two questions.

No account.

No cloud.

No exam engine.

No real vocabulary system.

No personal intelligence.

No internal content tools.

No serious recovery.

No idea that I would eventually be thinking about things like synchronization conflicts, evidence confidence, revision history or production security.

I just wanted a better way to practise verbal ability for IPMAT.

The first version needed to answer one question:

Would I actually use this?

I did.

That answer created about fifty more questions.

And somewhere along the way, the tiny practice app stopped being tiny.

So before I build another major piece of it, I wanted to stop and document what actually exists now.

Not what I imagine it becoming.

Not what is on some roadmap.

Not what would look good in a product announcement.

What works.

What partially works.

What still needs to be proved.

And what I absolutely should not pretend is finished.

The original loop is still there

Despite everything I have added, the basic interaction has not changed very much.

Open the app.

Practise something.

Understand what went wrong.

Come back later.

That still matters to me.

I do not want the project to become so impressed with its own systems that studying turns into operating software.

The additional machinery is supposed to make that loop better.

Not replace it.

The current practice system can build sessions around different topics and exam contexts.

Session size is no longer locked to the original six questions.

Timing can be relaxed, strict or absent depending on what I am trying to do.

Questions can be saved.

Mistakes can return.

Explanations can teach why the correct option works and why the others do not.

Response time can be recorded without forcing every learning session to feel like an exam.

The interface became more capable.

But the purpose is still extremely ordinary.

Practise verbal ability.

The question bank is still smaller than the machinery around it

This is an important distinction.

The internal content pipeline has become capable of handling much more material than the first version ever could.

There are draft items.

Review states.

Metadata.

Sources.

Revision history.

Publishing controls.

Multiple exam contexts.

A much larger content-production system behind everything.

But that does not mean every item inside that machinery is ready for a learner.

The current learner bank is still relatively small.

There are 64 original reviewed questions in the current build.

That is twice the first version.

It is also nowhere near where the content eventually needs to be for long-term serious use.

I am happy with that distinction.

A few weeks ago, I probably would have wanted the largest number possible.

Now I would rather say:

64 reviewed questions

than quietly convert thousands of drafts into:

Thousands of questions available.

Those are not the same claim.

The entire reason I built a content workflow was to stop pretending they were.

Vocabulary became its own product inside the product

The first vocabulary idea was almost embarrassingly simple.

Words.

Definitions.

Maybe some cards.

That version is basically gone.

The current bundled vocabulary collection contains 172 words and expressions.

More importantly, the system now treats vocabulary as something that needs repeated evidence.

Not:

I opened this card once.

Not even:

I answered one MCQ correctly.

A word can appear through different kinds of learning and testing.

Recognition.

Recall.

Synonyms.

Antonyms.

Context.

Typed retrieval.

Correct usage.

Confusable words.

Roots and families.

Idioms and phrasal verbs.

The current engine supports ten learning formats.

That sounds like a feature count.

I try not to think about it that way.

The useful part is that different formats expose different weaknesses.

I can recognise a word and still fail to recall it.

I can recall the meaning and still misuse it.

I can understand two similar words independently and confuse them when both appear together.

I can remember something today and forget it next week.

A single mastered / not mastered switch cannot describe that honestly.

So vocabulary now has memory.

Lapses.

Review timing.

Weakness signals.

Different kinds of evidence.

Recently forgotten items.

Words that are due.

Words that need more proof.

There is also a small placement diagnostic so the app does not have to assume every learner begins from zero.

The system is still heuristic.

It is not some scientifically perfect model of human memory.

But it is significantly more honest than the counter I started with.

The app now has a real exam mode

For a long time, exam mode existed mostly as:

Something I should build later.

Then later arrived.

And I learned very quickly why adding a countdown to normal practice would not have been enough.

Practice and examination behave differently.

When I am learning, the app should help me.

When I am sitting a mock, the rules should stop adapting themselves around my comfort.

The current exam engine understands multiple exam contexts.

IPMAT Indore.

IPMAT Rohtak.

JIPMAT.

IIM Bangalore UG.

IIM Kozhikode BMS.

Those configurations can carry their own timing and marking behaviour rather than forcing every exam into one generic preset.

A mock can have a fixed clock.

Positive and negative marking.

A question palette.

Marked-for-review state.

Answer persistence.

Automatic submission when appropriate.

And after the attempt, the result can show more than a final score.

Accuracy.

Attempted and unattempted questions.

Topic performance.

Time.

Speed against accuracy.

Answer review.

Patterns that may deserve attention.

This is probably the point where the project stopped feeling like a vocabulary side project even to me.

Then the app started interpreting what happened

Tracking is easy.

Interpretation is harder.

The app had already accumulated:

attempts,

mistakes,

timing,

exam results,

vocabulary history,

lapses,

activity,

and topic information.

I could have stopped there and built increasingly decorative dashboards.

Accuracy.

Hours.

Streaks.

Graphs.

Heatmaps.

Done.

Instead, I wanted the system to start asking:

What does this evidence suggest I should do next?

That became Personal Intelligence.

The current system can look for stronger and weaker topic signals.

It can consider how much evidence exists before becoming confident.

It can look at the relationship between speed and difficulty.

Compare recent periods.

Notice repeated errors.

Keep mistake notes and reasons.

Estimate memory risk.

Build targeted repair work.

What I care about most is not any individual insight.

It is the rule underneath all of them:

The app should know when it does not know enough.

Two wrong questions should not diagnose me.

One unusually fast answer should not declare a strength.

A tiny sample should remain a tiny sample.

I would rather the app say:

Not enough evidence yet

than manufacture a confident recommendation because the dashboard needs something to display.

Accounts changed what the app was allowed to get wrong

The earliest version lived inside one browser.

That was enough.

Now the application has authentication.

A learner can create an account.

Sign in.

Sign out.

Reset or update a password.

Use progress that can exist beyond one browser session.

The backend can preserve attempts, vocabulary state, saved items, exam sessions and mistake information.

That sounds like ordinary application infrastructure.

It also completely changed the responsibility of the project.

When everything belonged only to me, I could recover from strange behaviour manually.

Once information belongs to an account, the software needs stronger boundaries around whose information is whose.

The backend should enforce privacy instead of trusting the interface to behave correctly.

And once progress can move between devices, synchronization becomes something the user should not have to understand.

Local-first survived the move to the cloud

This was one of the decisions I cared about most.

Adding accounts did not mean I wanted every interaction to wait for the internet.

The app still records important progress locally.

If something needs to reach the server and cannot, it can wait in an outbox.

Reconnects can trigger pending work.

Duplicate events can be recognised instead of blindly recorded twice.

There are safeguards around conflicting state.

The app can create safety snapshots.

Progress can be exported and restored.

There is also a flow for deleting the account and synced information rather than making data permanent simply because it reached the cloud.

That is a lot more infrastructure than:

Save this answer.

But that complexity exists so the student-facing experience can remain simple.

Ideally:

I answer something.

The app remembers.

Everything else is somebody else's problem.

Unfortunately, I am currently somebody else.

Offline is better, but I am not calling it solved

This is one of the places where I want the checkpoint to be precise.

The app has strong local-first foundations.

That does not mean every offline scenario is finished.

There is still work around properly preserving and refreshing downloaded content.

A cold launch without connectivity needs stronger guarantees.

Reconnect behaviour needs more abuse testing.

Multiple tabs need testing.

Different devices need testing.

Account switching needs testing.

Stale state needs testing.

Interrupted writes need testing.

Storage problems need testing.

The architecture exists.

Parts of it work.

That is not the same thing as proving every ugly failure case.

So:

Local-first? Yes.

Completely proven offline product? No.

Not yet.

Active sessions became something worth recovering

This was another strange consequence of the app becoming useful.

Completed progress is one thing.

An unfinished session is another.

If I am halfway through practice and something disappears, what should survive?

If I am forty minutes into a mock, what does recovery mean without accidentally changing the conditions of the exam?

How should timing behave?

How do I avoid turning time spent away from the screen into learning evidence?

How do I avoid duplicating an attempt after a retry?

The app now treats active work more seriously than the first version ever needed to.

But this is also an area where implementation and proof are different.

Recovery logic can exist.

I still want more real-device failure testing before I call every interruption case dependable.

Apparently the correct way to test a study app eventually involves deliberately trying to ruin your own study session.

There is now another app behind the app

The Content Studio still makes me laugh slightly.

The project became complicated enough that editing content directly stopped being a reasonable long-term workflow.

So there is now a private system behind the learner application.

Questions.

Passages.

Vocabulary.

Sources.

Exam information.

Editorial states.

Imports.

Validation.

Review.

Revision history.

Rollback.

Publishing.

That internal tool exists because content is no longer allowed to become learner-facing just because it exists somewhere in a file.

Something can be a draft.

Then reviewed.

Then published.

Later corrected.

Or retired.

The distinction seems boring until one wrong explanation reaches a student.

Then it becomes very interesting.

The internal system also gives the app a much stronger foundation for growing the content bank without turning every update into manual surgery.

Reliability became a feature even though nobody can see it

Some of the work I am happiest with now creates almost nothing interesting to screenshot.

Duplicate prevention.

Snapshots.

Safer account boundaries.

Controlled error handling.

Production protections.

Internal tools that remain internal.

Cleaner release packaging.

Tests designed to catch unsafe assumptions.

None of that helps me make a dramatic before-and-after image.

It does change whether I trust the application.

The first prototype mostly needed to prove:

Can this happen?

The current project increasingly needs to prove:

Will this still behave correctly when things stop cooperating?

That is a much less glamorous question.

It is also much closer to software I would let somebody else use.

The app is not launch-ready

This is probably the most important sentence in the article.

The project has become substantial.

I am proud of that.

It is still not a finished public product.

The biggest problem is content.

Sixty-four reviewed questions can prove an engine.

They cannot sustain months of serious preparation.

The vocabulary bank demonstrates the learning system.

It is not yet the complete corpus I ultimately want to study from.

Content also needs stronger independent review before I would be comfortable putting large numbers of students through it.

Exam configurations need continued checking against current official information.

Reliability needs more real-world testing.

Offline behaviour needs more work.

There are edge cases I have not discovered yet.

There are definitely bugs I have not discovered yet.

That is normal.

What would be dishonest is pretending otherwise because the interface already looks polished.

There is no Android app yet

This is another thing worth stating clearly.

The current product is a responsive web application.

It works across the kinds of screens I use it on.

There is no finished Android package.

There is no finished Windows desktop application.

I have thought about both.

That does not make them real.

I could probably begin packaging the web application immediately.

I do not want platform count to become another vanity metric.

A native wrapper around unfinished behaviour would still contain unfinished behaviour.

The shared core needs to deserve more packaging first.

So web remains the actual product today.

Everything else can wait until there is a reason to make it somebody's installation problem.

The visuals are not the bottleneck anymore

This is another weird milestone.

Earlier versions constantly made me want to redesign things.

Now I mostly don't.

The interface is not perfect.

There will always be details to improve.

But the product does not currently need another dramatic visual reinvention.

It needs deeper content.

More testing.

More reliability.

Better evidence.

More actual use.

That is a nice problem to have.

The thing finally looks close enough to what I want that the hard work is increasingly somewhere underneath it.

So what is this now?

I have tried several descriptions.

Vocabulary app.

IPMAT practice app.

Verbal learning system.

Exam-preparation tool.

They are all partially correct.

The easiest description is probably still:

My IPMAT verbal app.

The difference is what those words now contain.

Question practice.

Vocabulary learning.

Review.

Progress.

Exam simulation.

Personal intelligence.

Accounts.

Synchronization.

Offline-first behaviour.

Recovery.

Content operations.

A lot of boring infrastructure.

And an increasing number of decisions designed around one principle:

Do not claim more than the evidence supports.

That rule started inside the learning engine.

It has slowly become how I think about the entire project.

The 32-question version did its job

I sometimes look back at the original prototype and see everything it lacked.

That is unfair.

It was never supposed to contain what exists now.

Its job was not to become a complete product.

Its job was to answer:

Would practising this way actually help me?

Yes.

That answer earned the next version.

Then the next version exposed another problem.

That problem earned another system.

And so on.

If I had tried to design the current application on day one, I probably would have spent months building architecture around assumptions I had never tested.

Starting badly was useful.

Starting small was useful.

Being the first user was useful.

The current complexity only makes sense because the simple version came first.

Where the app actually stands

If I had to describe it without marketing language:

It is a strong private alpha moving toward early beta.

The important systems exist.

A lot of them work well.

The learning architecture is far beyond the tiny tool I originally needed.

The product is useful to me today.

But the content is not deep enough.

Every reliability path has not been proven.

Every platform does not exist.

Every question has not received the level of review I ultimately want.

And I still have the luxury of breaking things privately before somebody else depends on them.

I want to use that luxury.

The goal is no longer to prove that I can build an IPMAT verbal app.

I did that a while ago.

The harder question now is:

Can I make it dependable enough that its intelligence, content and progress actually deserve a student's trust?

That is where the project stands.

Which is considerably further than 32 questions.

And considerably further from finished than the screenshots might suggest.