Back to all writing

I Taught My App the Difference Between Evidence and a Guess

IPMAT Series•Part 25•9 min read•By Mikhil

I Taught My App the Difference Between Evidence and a Guess

The easiest way to make software sound intelligent is to make it confident.

You struggle with Geometry.

Your retention is weak.

You are inefficient under pressure.

You should practise Algebra next.

Four sentences.

Very convincing.

Maybe completely wrong.

That became one of the biggest design problems in Intelligence.

Not:

How do I make the model generate more conclusions?

But:

How do I stop it from confusing a conclusion with an observation?

What happened is not the same as what it means

Suppose I answer a question incorrectly.

That happened.

The system can store it.

Now suppose it says:

Conceptual weakness.

That did not happen.

That is an interpretation.

Maybe I knew the concept and made a sign error.

Maybe I misread the question.

Maybe I changed from the correct answer to the wrong one.

Maybe I was rushing.

Maybe the question itself is bad.

Maybe I genuinely do not understand the concept.

Those possibilities should not be collapsed into one label just because the interface wants an insight.

So the learner model now makes a strong distinction:

Observation

versus

Inference.

Raw evidence should be boring

This is actually a good thing.

Raw evidence can be plain.

Attempt ID.

Question.

Skill.

Answer.

Correctness.

Timing.

Confidence.

Prior exposure.

Review event.

Mock event.

The boringness is useful because it keeps interpretation separate.

If the learner model changes later, the original evidence should still exist.

A new model can re-interpret the same history without rewriting what actually happened.

That is much safer than storing only the latest conclusion.

Inferences need versions

Suppose version 1 of the model says:

Likely execution bottleneck.

Then I improve the model.

Version 2 says:

Evidence is mixed; likely calculation instability instead.

I do not want the system pretending the second conclusion was always the truth.

The inference should have a version.

The model should know which rules produced it.

That creates an audit trail.

What did the system believe?

When?

Based on what?

With what confidence?

That may sound excessive for a study app.

It matters the moment the app starts influencing what somebody practises.

Confidence should be attached to claims

Not every insight deserves the same language.

Imagine:

One wrong answer.

Versus:

Twenty recent first-exposure questions across several difficulty levels showing the same pattern.

Those are not equivalent.

So the model carries evidence strength and uncertainty.

A weak pattern might deserve:

Possible issue. More evidence needed.

A stronger repeated pattern may justify:

Consistent weakness observed.

The wording should change with the evidence.

This is not cosmetic.

The confidence of the interface teaches the learner how seriously to take the conclusion.

Contradictory evidence should lower confidence

Models love clean stories.

Humans are messy.

Maybe my Algebra accuracy is strong.

My timed Algebra is weak.

My recent work is improving.

My older history is poor.

My confidence is low even when answers are correct.

What am I?

Strong?

Weak?

Improving?

Uncertain?

Potentially all of those statements describe part of the evidence.

The system should not delete inconvenient signals to create one neat label.

Contradictions are information.

They should lower confidence or make the explanation more specific.

The learner model can be wrong

I wanted this assumption built into the architecture.

Not added later as a disclaimer.

The model is a hypothesis engine.

It is trying to explain incomplete evidence.

That means being wrong is not an edge case.

It is expected.

The question is whether the product has mechanisms to notice, revise and recover.

New evidence should be able to weaken an old conclusion.

Interventions should be able to fail.

Recommendations should be dismissible.

The system should be capable of changing its mind.

That feels much more intelligent than permanent certainty.

Recommendations are hypotheses too

Suppose Intelligence recommends:

Do a timed Arithmetic repair set.

The hidden claim is:

This intervention is likely to help.

That can be tested.

Did the learner accept it?

Did they complete it?

What happened afterward?

Did the relevant weakness improve?

Did nothing change?

Did another problem become clearer?

Now the recommendation has an outcome.

That outcome can become evidence for future recommendations.

The system starts learning not only:

What is wrong?

but:

What kind of intervention tends to help?

Success conditions make recommendations measurable

I do not want vague advice.

Practise more Quant.

That is technically advice.

It is almost useless.

A stronger recommendation might be:

Repair the unresolved percentage errors.

Then complete a fresh timed set.

Success means accuracy remains stable while median active time improves without increased guessing.

Now the recommendation can be evaluated.

Did the learner meet the condition?

If yes, great.

If no, maybe the original diagnosis was incomplete.

That creates a proper loop.

Readiness is one of the most dangerous inferences

People care about readiness.

That makes it tempting to turn it into a dramatic score.

The model deliberately avoids treating readiness like a prophecy.

It needs enough relevant evidence.

Recent evidence.

Timed evidence.

Coverage.

Mock behaviour.

Confidence.

Contradictions.

If those conditions are weak, the system should suppress the conclusion.

Not enough evidence is a valid readiness state.

That is much more useful than a fake percentage with two decimal places.

Cross-subject reasoning makes uncertainty more important

Once Intelligence sees Verbal and Quant, the model gets more powerful.

It also gets more opportunities to make nonsense.

Maybe Verbal is improving.

Quant is not.

Does that mean study allocation is wrong?

Maybe.

Maybe Quant questions were harder.

Maybe the student recently focused on Verbal because it was weaker.

Maybe Quant evidence is old.

Cross-subject conclusions need context too.

Connecting more data does not remove uncertainty.

Sometimes it multiplies it.

The model should explain the evidence chain

If the app says:

Timed execution appears to be limiting your Quant performance

I want to be able to inspect why.

Recent timed sets.

High conceptual accuracy.

Slow median pace.

Repeated overinvestment.

Stable untimed performance.

That chain matters.

The learner should be able to distinguish:

This conclusion has evidence

from:

The app generated a sentence.

Explainability is not about exposing every mathematical detail.

It is about showing enough of the reasoning that the claim can be challenged.

Evidence coverage is its own signal

Sometimes the problem is not weakness.

It is absence.

Maybe I have strong Arithmetic evidence and almost no Geometry evidence.

A normal dashboard may show Geometry at 0%.

That is misleading.

Zero performance and zero evidence are not the same thing.

So coverage itself matters.

The model needs to know where it has enough observations to say something.

A gap in the evidence should not automatically become a gap in ability.

That seems obvious.

It is easy to get wrong in software.

Synthetic development data has to stay labelled

Intelligence can run in prototype mode using synthetic raw events.

That is useful because the same pipeline can be exercised before full integration.

But synthetic evidence is not validation.

A beautiful insight generated from demo events proves the engine can process events.

It does not prove the model describes real learners accurately.

That boundary needs to stay visible.

Otherwise development convenience turns into accidental product claims.

Server-authoritative derived state became important

If official learner-state conclusions can be written directly by the browser, trust becomes messy.

So the shared platform separates raw client evidence from official derived Intelligence.

The server can verify the learner, load the relevant evidence and persist versioned derived state.

That reduces the chance that one client accidentally rewrites the learner model with stale or malformed conclusions.

Again:

Boring architecture.

Important consequence.

Append-only evidence protects history

Normal clients should not casually rewrite old attempts.

If an answer happened, it happened.

Later notes can change.

Interpretations can change.

The original evidence should remain.

That is why append-only patterns became useful for raw learning events and derived-state history.

The system can build a timeline instead of repeatedly editing the past into whatever the current model believes.

This creates a strange kind of humility

Software usually wants to feel finished.

A learner model built around uncertainty has to admit things like:

Possible.

Likely.

Low confidence.

Contradictory evidence.

Not enough evidence.

I used to think those labels would make the product feel weaker.

Now I think they make it feel more serious.

Because the alternative is confidence theatre.

AI is not exempt from this

The same principle applies anywhere I use AI-assisted evaluation.

A language model can produce useful structured feedback.

That does not make its judgment objective truth.

Subjective evaluation needs confidence.

Fallbacks.

Reviewed examples.

Disagreement handling.

And it should not be allowed to single-handedly promote someone to "mastered" because one generated sentence looked good.

The app should treat AI output as another kind of evidence.

Not an oracle.

Intelligence should become more useful as it becomes less dramatic

I do not want:

AI HAS DISCOVERED YOUR SECRET LEARNING TYPE.

I want:

You have enough evidence to say this.

You do not have enough evidence to say that.

This pattern repeated.

This recommendation worked before.

This one did not.

This area is improving.

This area is uncertain.

Here is the next useful action.

That feels less magical.

It is much more valuable.

The difference between smart and persuasive

A persuasive system can tell a clean story from messy evidence.

A smart system should know when the clean story is unjustified.

That is what I am trying to build.

Not a dashboard that always has an answer.

A learner model that understands the difference between:

I observed this.

I infer this.

I am confident about this.

and

I do not know yet.

That may be the most important intelligence feature in the entire ecosystem.

Not another model.

Not another chart.

The ability to stop before a guess becomes a fact.