Back to all writing

How Do You Build 2,500 Good Exam Questions Without Filling the App With Garbage?

IPMAT Series•Part 11•11 min read•By Mikhil

How Do You Build 2,500 Good Exam Questions Without Filling the App With Garbage?

At this point, the app had a strange problem.

The systems were becoming stronger than the content.

I had built a vocabulary engine.

An exam engine.

Personal intelligence.

Offline reliability.

Sync protection.

Account controls.

The app could do far more with a good question than the first prototype ever could.

Which meant the obvious next move was:

Add a lot more questions.

That sounds simple.

It is not.

Because there are two very different ways to build a large question bank.

One is:

Make the number big.

The other is:

Make the questions useful.

I wanted the second one.

Big numbers are easy to fake

This is especially true now.

If I wanted to, I could generate thousands of questions very quickly.

Vocabulary MCQs.

Grammar questions.

Reading passages.

Sentence correction.

Synonyms.

Antonyms.

Context usage.

Enough content to put a huge number on the landing page.

10,000+ questions

Looks impressive.

But that number would not tell me:

How many are actually relevant.

How many are duplicates in disguise.

How many have weak distractors.

How many have ambiguous answers.

How many are too easy.

How many test the wrong skill.

How many have bad explanations.

How many accidentally teach something incorrect.

How many resemble the exams I am actually preparing for.

A large question bank can be extremely bad while looking extremely complete.

That is what I wanted to avoid.

The app was no longer only about IPMAT Indore

Another complication was scope.

The product had gradually become useful across a wider group of IIM undergraduate entrance exams.

That meant I was now thinking about:

IPMAT Indore

IPMAT Rohtak

JIPMAT

IIM Bangalore UG

IIM Kozhikode BMS

These exams overlap.

But they are not identical.

Treating everything as one generic "IPMAT question bank" would be convenient.

It would also flatten differences that actually matter.

So the content system needed to understand exam context too.

One question can be useful in multiple places

This led to an architectural idea I really liked.

A question itself can be canonical.

The same underlying item may be relevant to more than one exam.

But the way the system uses that question can depend on the exam context.

That means I do not need five physically separate copies of the same useful question.

Instead, one question can have multiple exam-specific views or mappings.

That keeps the content cleaner.

It also makes it easier to reason about where a question belongs.

The important distinction becomes:

What is this question?

and

Where is this question relevant?

Those should not be the same piece of information.

Research had to come before scale

This brought me back to one of the earliest lessons from the project.

Do not begin with the database.

Begin with the exam.

Before generating at scale, I needed a better evidence map.

What skills are actually tested?

What kinds of verbal questions recur?

Which topics appear across multiple exams?

Which ones are more specific?

What level of difficulty makes sense?

What should count as core content?

What should be lower priority?

That work is less exciting than generating questions.

It is also what stops the generator from producing thousands of beautifully formatted irrelevant items.

The research pack became much larger

By this stage, I had built up a much broader research base.

Previous-year patterns.

Exam structures.

Preparation material.

Topic evidence.

Different verbal categories.

The goal was not to copy questions.

The goal was to understand the space well enough to create original material that belongs there.

That distinction matters.

I do not want a product built by taking existing books, changing three words and pretending the result is original.

I want the research to inform what I build.

Not become the thing I ship.

Then the numbers got weird

As the content architecture developed, the draft system reached roughly 2,190 canonical questions.

Because many of those questions could be relevant in more than one exam context, they represented around 5,000 exam-specific views.

Those numbers sound huge compared with the first version.

Thirty-two questions.

Then 64.

Now thousands of potential question placements.

But I had to be very careful about what those numbers meant.

They did not mean:

2,190 perfect production-ready questions.

They meant:

2,190 canonical draft items inside the content pipeline.

That is a very different claim.

Draft is not published

This distinction became one of the most important parts of the content system.

A question can exist without being ready for students.

That sounds obvious.

But if the database treats every generated item as immediately usable, quality collapses quickly.

So I wanted a clear separation between stages.

Research.

Draft.

Review.

Approved.

Published.

The exact workflow can evolve.

The principle should not.

Creation and publication are different actions.

That gives me room to generate at scale without pretending scale equals quality.

Stable IDs matter here too

Once the question bank becomes large, identity becomes important.

A question should not become a completely new object every time I revise the wording.

The system needs stable identifiers.

That helps with:

Review history.

Exam mappings.

Versioning.

Analytics.

Replacing bad content.

Tracking what changed.

Preventing accidental duplicates.

This is another example of something that feels unnecessarily formal at 32 questions and completely necessary at 2,000.

Scale changes what counts as overengineering.

The distractors are often harder than the answer

This is one of the biggest content-quality problems.

Writing a correct answer is easy.

Writing three wrong answers that are:

Plausible.

Clearly wrong.

Not accidentally correct.

Not obviously stupid.

Similar enough to require thought.

But not so ambiguous that two options could work.

That is much harder.

A weak distractor turns a question into a recognition game.

The student does not need to know the answer.

They only need to notice that the other three options are ridiculous.

That makes the question look functional while testing almost nothing.

So reviewing distractors became just as important as reviewing the answer itself.

Explanations are part of the question

I also stopped thinking of explanations as optional text added after the real content.

For a learning product, the explanation is part of the item.

A good explanation should tell me:

Why the correct option works.

Why the tempting alternatives do not.

What concept matters.

What I should notice next time.

A question without a useful explanation may still work inside a strict mock.

But it is much less valuable inside practice and repair.

And because the same content can feed multiple engines, explanation quality matters more now than it did in the first prototype.

Difficulty cannot just be vibes

Another problem is difficulty.

It is very easy to label something:

Easy.

Medium.

Hard.

Based on intuition.

That becomes unreliable at scale.

A vocabulary word may feel difficult to me but obvious to someone else.

A grammar question may look simple but contain a very deceptive distinction.

A long question is not automatically hard.

An obscure word is not automatically useful.

So difficulty needs to be treated as something that can become better calibrated over time.

Initial labels can exist.

But real attempt data should eventually help refine them.

That creates a nice feedback loop.

The content bank teaches students.

Student performance can eventually teach the content system something too.

AI is useful here, but dangerous

AI makes this entire pipeline much more possible.

There is no realistic way I personally write thousands of high-quality questions manually in a short period while also preparing for IPMAT and building the product.

AI can help generate drafts.

Create variations.

Produce distractors.

Suggest explanations.

Transform research into candidate questions.

That is incredibly useful.

But it creates a dangerous temptation:

Generate → insert into database → done.

I do not trust that workflow.

AI can produce questions that look perfectly professional while containing subtle problems.

An ambiguous answer.

A false explanation.

A bad synonym.

A grammatical edge case.

A question testing something irrelevant.

Fluent text is not the same thing as reliable educational content.

So AI belongs inside the pipeline.

Not above it.

We needed a manageable production rhythm

At one point, I considered building the entire question bank before doing anything else.

That would mean stopping feature development and spending a long stretch almost entirely on content.

I decided against it.

The better approach for me is batches.

Roughly 250 questions at a time.

Generate.

Review.

Fix.

Integrate.

Test.

Then move to the next batch.

That number is large enough to make visible progress.

But small enough that quality control does not become completely overwhelming.

It also gives me opportunities to improve the pipeline between batches.

If batch one exposes a problem, I can fix the process before creating another thousand questions with the same flaw.

The first serious target is 2,500

I eventually settled on around 2,500 strong questions as the first major content target.

Not because 2,500 is some magical number.

It is simply enough to make the product substantially more useful.

Enough variety for practice.

Enough depth for repair.

Enough content for mocks to improve.

Enough material for personal intelligence to have more evidence.

Enough that the app begins feeling like a serious preparation tool rather than a technical demo.

And importantly:

2,500 is not the finish line.

It is a launch-scale target.

The bank can continue growing afterward.

This changed how I think about launch

Earlier, I unconsciously imagined launch as:

Finish everything.

Then release.

But this project probably does not need a final content state.

There will always be more useful questions to add.

More vocabulary.

More exam evidence.

More edge cases.

More refinements.

So the real question is:

When does the product contain enough high-quality material to become genuinely useful?

That is a much healthier threshold.

It means I can launch something strong and continue improving the bank afterward.

Quality control became its own product problem

The more I thought about this, the more I realised content tooling itself matters.

If I eventually have thousands of questions, I need efficient ways to inspect them.

Filter them.

Find duplicates.

Check exam mappings.

Review explanations.

Approve batches.

Reject weak items.

Correct metadata.

Track publication state.

At some point, the content pipeline becomes almost a second application behind the student-facing product.

Students should never see that complexity.

But somebody has to manage it.

Right now, that somebody is me.

This is probably the least visible hard part

People looking at the app will see:

The dashboard.

Vocabulary.

Mocks.

Progress charts.

Animations.

Streaks.

The question screen.

They probably will not think about the process required to make sure the next question is worth answering.

But educational software lives or dies on content.

A perfect interface serving bad questions is still bad exam preparation.

A sophisticated adaptive engine making decisions from weak material becomes sophisticated garbage distribution.

The intelligence only matters if the content underneath it deserves to be selected.

The app is ready for the content now

This is what feels different from earlier.

When I had 32 questions, I needed more questions because I kept seeing the same ones.

Now I need more questions because the systems around them are ready to use them properly.

A vocabulary question can influence mastery.

A practice question can reveal a weakness.

An exam question can affect mock strategy and timing evidence.

A repeated mistake can trigger repair.

A question can contribute to personal intelligence.

The content now has somewhere meaningful to go.

That makes scaling it feel much more worthwhile.

I am glad I did not start with 10,000 questions

If I had generated a giant bank at the beginning, I probably would have designed the system around whatever data structure happened to exist.

Instead, the product evolved first.

Now I know much more clearly what each question needs to support.

Stable identity.

Skill tags.

Difficulty.

Exam relevance.

Explanations.

Review state.

Publishing state.

Enough metadata for the engines around it to work.

The delay probably saved me from rebuilding thousands of items later.

So now comes the repetitive part

Build 250.

Review.

Integrate.

Build another 250.

Review.

Integrate.

Keep going until the bank reaches the first serious target.

Not glamorous.

But honestly, I like that the project has reached this stage.

The hard questions are becoming less like:

Can I build this?

and more like:

Can I make this good enough to trust?

That feels like progress.

The app started because I had a vocabulary backlog.

Then it became a practice tool.

Then a learning system.

Then an exam engine.

Then something that could interpret my preparation.

Now the underlying machinery is finally strong enough that the next challenge is simply giving it enough good material to work with.

Around 2,500 questions should be a decent start.