Back to all writing

I Put AI in the App—and Refused to Let It Decide What I Know

IPMAT Series•Part 14•12 min read•By Mikhil

I Put AI in the App—and Refused to Let It Decide What I Know

Type a sentence using:

Obdurate.

That sounds like a much better vocabulary test than choosing one definition from four options.

Until you actually try building it.

Suppose I write:

He remained obdurate despite everyone asking him to reconsider.

Good.

Now suppose I write:

The obdurate weather ruined our picnic.

Probably not good.

What about:

His obdurate attitude made the discussion difficult.

Also reasonable.

And then there are sentences that are grammatically correct but reveal that I only vaguely understand the word.

Or sentences where the meaning is right but the wording is awkward.

Or sentences that technically work in some unusual context.

A multiple-choice question has a known answer.

An open-ended sentence does not.

That was the problem.

And for the first time in this project, AI started looking genuinely useful.

I had deliberately avoided putting AI everywhere

This project has been built with a lot of AI assistance.

That is not a secret.

AI helps me write code.

Debug things.

Think through architecture.

Question decisions.

Find edge cases.

Turn ideas into working implementations much faster than I could alone.

But that never meant the product itself needed an AI button on every screen.

Actually, I wanted to avoid that.

There is a strange pressure around software now where adding AI automatically makes something sound more advanced.

AI-powered learning.

AI tutor.

Personalised AI coach.

AI-generated insights.

Put a sparkle icon beside it and apparently the product has evolved.

I did not want to do that.

If a deterministic system could solve something reliably, I preferred the deterministic system.

If an answer was already known, I did not need a language model improvising around it.

If a rule could be explained clearly, I did not need AI pretending the rule was its opinion.

The vocabulary engine had already become adaptive without needing a model to decide everything.

It could track evidence.

Remember lapses.

Schedule reviews.

Identify weak areas.

Choose between learning modes.

Distinguish different kinds of performance.

None of that required me to hand the entire learning process to an AI system.

So for a long time, AI stayed outside the student-facing experience.

Then sentence practice created a problem that ordinary rules were much worse at solving.

Multiple choice could test recognition

Typed recall could test retrieval.

Context questions could test interpretation.

Confusable-word practice could test discrimination.

But eventually I wanted to ask something harder:

Can you actually use the word?

That requires production.

Not recognising an answer.

Not selecting a sentence somebody else wrote.

Producing language yourself.

So I added sentence practice.

The basic interaction is simple.

The app gives me a word.

I write a sentence using it.

Then I need some way to understand whether that attempt was useful.

And immediately, the neat deterministic world starts falling apart.

There is no single answer key for a sentence

With an MCQ, I know exactly which option is correct.

With typed vocabulary recall, I can compare against expected meanings and accept a reasonable amount of variation.

Sentence usage is different.

There may be hundreds of completely valid answers.

The same word can appear in different grammatical structures.

Different tones.

Different contexts.

Different sentence lengths.

A rigid checker would either reject too much or accept things it should not.

I could have created pattern matching.

Required certain words nearby.

Checked grammar mechanically.

Built increasingly complicated rules.

But at some point I would just be creating a worse language model out of if-statements.

This was one of the first places where the probabilistic nature of AI actually matched the problem.

Language is flexible.

The evaluation needs some flexibility too.

That did not mean I wanted AI in charge.

First, I made sentence practice work without AI

This part mattered to me.

I did not want the feature to become unusable just because an AI service was unavailable.

The rest of the application had already pushed me toward local-first thinking.

Study should not suddenly become impossible because the network disappears.

And one learning mode should not require an external system merely to exist.

So sentence practice needed a useful baseline without AI.

The student can still see the word.

Think.

Write.

Compare their attempt against the meaning and example usage.

Reflect on whether they actually used it correctly.

The exercise itself already has value because production is harder than recognition.

AI feedback can improve that experience.

It should not create the experience from nothing.

That distinction also gave me a clean rule:

AI is an enhancement here, not a dependency.

Then I added optional AI feedback

This is where the feature became more interesting.

When AI feedback is available and deliberately enabled, the student's sentence can be evaluated more flexibly.

Not with some dramatic:

CORRECT / INCORRECT

judgment.

More like feedback.

Does the sentence appear to use the intended meaning?

Is there a likely misuse?

Is the grammar making the meaning unclear?

Could the sentence be improved?

What might the student be misunderstanding?

That is a much better use of a language model than asking it to control the entire learning system.

It is doing something language models are naturally suited to:

responding to language.

But I didn't want the AI secretly deciding mastery

This became the most important boundary.

Imagine I write a sentence.

The model says:

Excellent usage!

The easiest implementation would be:

Great.

Increase mastery.

Move the word forward.

Count the evidence as correct.

Maybe even mark the word learned.

I did not want that.

Because now a probabilistic model has quietly become the authority over the student's learning state.

And language models can be wrong.

They can misunderstand context.

They can accept weak usage.

They can reject unusual but valid usage.

They can sound incredibly confident while doing either.

That means AI feedback and mastery evidence should not automatically be the same thing.

The AI can help me understand my attempt.

It should not get to rewrite the app's belief about what I know simply because it produced an enthusiastic paragraph.

So I drew a boundary.

AI feedback cannot award mastery.

That might sound overly cautious.

I think it is necessary.

Advice and evidence are different things

This project keeps forcing me to separate concepts that initially look identical.

Seeing a word is not knowing it.

Analytics are not intelligence.

Attempting an action is not verifying it.

A draft is not published content.

And now:

Feedback is not evidence.

AI can say:

This sentence appears to use the word correctly.

That can be genuinely useful.

But the structured learning system should still build confidence from controlled evidence.

Recall.

Recognition.

Retention.

Repeated performance.

Context.

Lapses.

Other interactions where the rules are clearer.

The AI can sit beside that system.

It does not need to become that system.

This also makes the product easier to reason about.

If my vocabulary state changes, I want to understand why.

Not discover that some invisible language-model judgment moved a score somewhere.

I made the AI part explicit

There was another thing I did not want.

Sending someone's writing somewhere without making that obvious.

A user typing a sentence into a local exercise and a user deliberately requesting AI feedback are not the same interaction.

So the AI layer needed to be optional.

Not quietly running behind every keystroke.

Not pretending to be ordinary local validation.

Not collecting language just because it might be useful.

If AI feedback is used, the student should know they are asking for AI feedback.

That sounds basic.

I think it should be basic.

The more invisible AI becomes inside software, the easier it is for users to stop understanding what is actually happening to their input.

For this feature, I prefer explicitness.

This is also why I don't call it an AI tutor

That label would be easy.

And impressive.

And currently dishonest.

The system can provide bounded feedback on a particular exercise.

That does not make it a tutor.

A tutor understands the larger lesson.

Recognises misunderstanding over time.

Chooses how to explain something differently.

Knows when to intervene.

Knows when to leave the student alone.

Connects one concept to another.

Understands what has already been taught.

Adapts its teaching intentionally.

Maybe software can eventually do more of those things.

This feature does not prove that.

It is an AI-assisted feedback layer.

That description is much less exciting.

It is also much more accurate.

The structured system still does most of the work

This was probably the most reassuring thing about adding AI.

The app did not suddenly need to be rebuilt around it.

Everything underneath remained useful.

The vocabulary library.

Mastery states.

Review scheduling.

Confusable words.

Roots and families.

Context practice.

Placement.

Daily missions.

Retention signals.

Lapses.

Different learning modes.

All of that still exists.

AI simply gained a narrow place where flexible language understanding is useful.

I like that architecture much more than:

Send everything to a model and hope it figures out what to do.

The structured systems create predictability.

AI adds flexibility where predictability becomes difficult.

They solve different problems.

I also stopped treating every successful interaction as equal

While working on this, I fixed another subtle issue in the learning logic.

Imagine I show a student the definition of a word.

Immediately afterward, I ask a question about that same word.

They answer correctly.

Did they remember it?

Technically, yes.

For about thirty seconds.

That is not the same evidence as recalling it tomorrow without assistance.

The earlier system could accidentally treat those interactions too generously.

So assisted answers immediately after learning needed to be understood for what they were:

practice.

Useful practice.

But not strong proof of lasting recall.

That distinction became even more important once the vocabulary system contained many different learning interactions.

A correct answer only means something when I understand the conditions under which it happened.

Was the definition visible?

Had the word just been introduced?

Was the answer recognised or recalled?

Was there a hint?

Was this a delayed review?

Context changes the meaning of the evidence.

The more sophisticated the engine becomes, the less comfortable I am with pretending all green ticks are equal.

AI created a new kind of trust problem

Normal software can fail very obviously.

A button does nothing.

A page crashes.

A request returns an error.

AI can fail beautifully.

That is more dangerous.

It can produce a polished explanation.

Confident tone.

Perfect grammar.

Very convincing structure.

And still be wrong.

That means the interface cannot rely on presentation to communicate certainty.

If something is AI-generated, I want that fact to remain understandable.

And I do not want one generated response overruling the structured evidence the rest of the app has collected.

This is similar to something I learned while building SYNTHESIS.

Intelligence and authority are different things.

A component can be capable of making a suggestion without being allowed to control the entire system.

The same principle works here.

The model can advise.

The learning engine remains responsible for the student's persistent state.

This was the first AI feature that earned its place

That is probably why I actually like this feature.

It did not begin with:

How can I add AI?

It began with:

How do I evaluate an answer that can be valid in hundreds of different forms?

Then AI became one possible answer.

That order matters.

There is a huge difference between finding a problem for a technology and finding a technology for a problem.

I have done enough feature-building now to know which direction usually produces worse products.

If I started with:

The app needs AI

I could have added a chatbot months ago.

A floating assistant.

A "Generate explanation" button.

AI-made study plans.

AI-generated motivational messages.

An AI avatar telling me to maintain my streak.

Plenty of things would look impressive in screenshots.

I did not need them.

Open-ended language feedback was different.

There was an actual gap.

AI made the interaction more capable.

So it stayed.

I still don't trust it enough to disappear into the background

Maybe that changes someday.

Maybe models become reliable enough in narrow educational tasks that some of these boundaries move.

Right now, I prefer caution.

AI feedback remains optional.

The core learning experience should survive without it.

Its output should be treated as assistance.

It should not silently create mastery.

And the student should understand when they are interacting with it.

Those rules make the feature slightly less magical.

Good.

I am increasingly suspicious of software that becomes impressive by hiding how much uncertainty sits underneath it.

The funny part is that AI made the app less "AI-powered" in my head

Before implementing this, I imagined adding AI as some major transformation.

After implementing it, I think about it almost like another tool.

A useful one.

Powerful in the right place.

Unnecessary in many others.

The application does not become intelligent merely because one screen can call a language model.

Most of the intelligence I care about still comes from the structure underneath.

Remembering what happened.

Knowing how much evidence exists.

Understanding the difference between new and forgotten.

Separating recognition from recall.

Tracking time.

Detecting patterns.

Choosing useful review.

Knowing when confidence is not justified.

AI can help around the edges of that system.

It does not replace the system.

I finally put AI inside the app

Just not in the way I would have expected when I started.

There is no giant chatbot waiting on the home screen.

No magical tutor claiming to understand everything about me.

No model deciding whether I have mastered a word.

Instead, there is one narrow problem where language is messy enough that flexible feedback helps.

I write a sentence.

I can ask for feedback.

The AI can respond.

I can learn from it.

And then the structured learning system continues doing what it was already designed to do.

That feels much healthier to me.

AI should make the product better where it has earned the responsibility.

Not become the product because the technology happens to be exciting.

For the first time, I think it earned a small place.

So I gave it one.

And then I put boundaries around it.