It Worked for Me. That Wasn’t Good Enough Anymore
For a long time, I had an extremely convenient quality standard for this app.
Does it work for me?
If something broke, I knew what I had been doing.
I could open the developer tools.
Restart the app.
Clear some state.
Fix the code.
Run it again.
If a weird edge case destroyed some test progress, that was annoying.
But I was also the developer, the tester and the only person depending on it.
That changes once an application starts holding things that matter.
Accounts.
Practice history.
Vocabulary state.
Exam attempts.
Bookmarks.
Mistake notes.
Recommendations.
Synced progress.
Suddenly:
“It usually works on my laptop”
starts sounding much less impressive.
The project had reached the point where I needed to stop treating reliability as something I could personally work around.
The software needed guardrails.
The app had become capable enough to hurt itself
The strange thing about adding features is that every useful capability creates another way something can go wrong.
Local progress is convenient.
Now local state can become stale.
Cloud sync is useful.
Now local and cloud state can disagree.
Accounts are useful.
Now data must belong to the correct account.
Offline support is useful.
Now changes can exist without immediately reaching the server.
Exam sessions are useful.
Now interrupted state matters.
An internal Content Studio is useful.
Now I need to make sure ordinary users cannot somehow reach internal functionality.
External services are useful.
Now failures outside my application need to be handled without turning into nonsense inside it.
None of these problems existed in the 32-question version.
That app barely knew anything.
There was not much to protect.
The current version knows significantly more.
That means there is significantly more it can get wrong.
Accounts made mistakes more serious
Adding authentication initially felt like a normal product milestone.
Create an account.
Sign in.
Keep progress.
Use another device.
Reset a password.
Sign out.
Pretty standard.
But accounts introduce a much more important idea:
This information belongs to someone.
That means the application cannot simply treat all stored state as one large bucket.
My attempts should be mine.
My saved questions should be mine.
My vocabulary history should be mine.
My exam sessions should be mine.
And signing into a different account should not create some strange mixture of two people's histories because the browser happened to remember both.
Once I thought about it that way, authentication stopped feeling like a login screen.
The login screen is the easy part.
The hard part is everything that has to remain true after the user is logged in.
Signing out is also a feature
This sounds ridiculously obvious.
If an app has:
Sign in
then surely it has:
Sign out.
But the deeper question is:
What does signing out actually mean?
Does local progress remain visible?
Does pending sync state belong to the old account?
What happens on another device?
What if I want every session signed out?
What happens after a password change?
Can I recover access if I forget the password?
Those are all much less glamorous than creating the original authentication form.
They are also part of whether an account system deserves to be trusted.
So the account layer grew beyond simply:
Email + password = user.
It now includes the ordinary boring things I would expect from software I actually use.
Sign-up.
Sign-in.
Sign-out.
Global sign-out.
Password reset.
Password update.
None of those features deserves a cinematic product demo.
I would be suspicious if they were missing.
The cloud should not become one giant shared notebook
The backend also needed actual boundaries.
It is not enough for the interface to hide somebody else's information.
The data layer itself should understand that one learner cannot simply request another learner's private progress.
That is a very different kind of protection.
A hidden button is UI.
An actual access rule is a boundary.
This distinction became increasingly important as more personal state moved into the backend.
Attempts.
Word reviews.
Saved items.
Exam sessions.
Mistake notes.
Preferences.
All of those are useful because the app can preserve continuity.
They also create responsibility.
A learning application does not need to be creepy to mishandle privacy.
It only needs one lazy assumption about who should be allowed to read what.
I did not want privacy to depend on the interface behaving politely.
Sync created a completely different class of bugs
One device is wonderfully simple.
There is only one version of reality.
Then you add another device.
Now imagine this:
I save a question on my laptop.
My phone still has an older local copy.
The phone goes offline.
I change something there too.
Later both reconnect.
Which state wins?
Or I delete something on one device.
Another device still remembers the old version.
Should the deleted thing magically come back?
Or the same attempt reaches the server twice because a retry happened at exactly the wrong moment.
The user did one thing.
The history now says they did two.
These are not exciting problems.
They are very good at destroying trust.
That is why the sync system eventually needed concepts such as duplicate prevention, deletion records that survive synchronization, reconnect-triggered work and an outbox for changes that have not safely reached the cloud yet.
The goal is not to make sync look intelligent.
The goal is to make it boring enough that I stop thinking about it.
A retry should not mean “do it twice”
This became one of those tiny rules that affects much more than it sounds like.
Network requests fail.
So software retries them.
Great.
Except some operations should happen exactly once.
Imagine an attempt being recorded.
The connection becomes uncertain.
The app does not know whether the server received it.
So it tries again.
Now suppose the server actually received both requests.
Congratulations.
One practice attempt has become two pieces of evidence.
That can affect activity.
Accuracy.
Recommendations.
Progress.
Maybe even what the app thinks I am weak at.
A duplicate is not merely an ugly database row anymore.
It can distort the learning system.
So important events needed identity.
The system needed a way to recognise:
I have already processed this.
That is one of those things nobody sees when it works.
Exactly how I like it.
I added safety snapshots because sync is not a backup
This distinction took me longer to appreciate.
If two devices perfectly synchronize the same bad state, synchronization has succeeded.
That does not mean the data is safe.
Imagine something goes wrong.
A bug overwrites useful progress.
Sync faithfully distributes that mistake.
Excellent synchronization.
Terrible outcome.
So I wanted another layer of protection.
The app now creates safety snapshots around important progress state.
There is also a portable export and restore path.
That gives the student another way to preserve and recover their history instead of assuming:
The cloud exists, therefore nothing can ever go wrong.
Cloud storage is useful.
It is not magic.
Deleting an account had to actually mean something
Account creation is fun product work.
Deletion is much less exciting.
But if an application lets somebody create an identity and store personal learning history, they should also have a way to leave.
Not:
Email me and maybe I will manually delete something later.
An actual deletion flow.
That required thinking about both the account and the synced information attached to it.
Again, this is not a feature that makes the homepage more impressive.
It is part of giving the user control over information the application only has because they trusted it in the first place.
The same thinking affected analytics.
I did not want the product quietly collecting detailed behavioural information merely because analytics tools make that easy.
Anonymous analytics are disabled by default in the current build, and the events I am willing to consider are deliberately limited.
The question should not be:
What can I collect?
It should be:
What do I genuinely need?
Then I started auditing the application like I didn't trust myself
This was probably the healthiest change.
When you build your own software, you naturally understand what it was supposed to do.
That makes it incredibly easy to test intention instead of reality.
Of course the internal editor is private.
I built it to be private.
Of course configuration files will never accidentally end up somewhere they should not.
I know they should not.
Of course an external service will return the format I expect.
That is what it normally returns.
Those sentences are dangerous.
So I started looking at the project more like something I had been handed by somebody else.
What assumptions would I question?
What would I try to break?
What should never appear inside a release archive?
Which internal pages should not exist in a public production build?
What happens if an API returns HTML instead of JSON?
What happens if a request is enormous?
What happens if authentication information is malformed?
What happens if an external service hangs?
What information reaches the user when something fails?
The more I looked, the more obvious it became that “works” was not a useful enough category.
Errors should fail like errors
One bug made this especially obvious.
A part of the app expected JSON from a request.
The route it called did not exist correctly.
The server returned an HTML error page.
The client happily tried to parse that HTML as JSON.
The result was the kind of error message developers immediately recognise:
Unexpected token '<'
That is technically understandable.
It is also terrible product behaviour.
The useful lesson was not this specific bug.
It was the assumption underneath it.
I had written code that effectively said:
The response will be what I expect.
Product-grade code has to ask:
What if it isn't?
So the hardened version became more defensive.
Check what came back.
Bound what can be sent.
Validate important identifiers.
Do not wait forever for an external request.
Return controlled errors instead of leaking whatever internal system happened to fail.
That is much less exciting than building another learning mode.
It is also exactly the kind of work that makes everything above it safer.
Internal tools should stay internal
The Content Studio created another interesting problem.
I built an entire private interface for managing questions, vocabulary, sources and editorial state.
The actions were already protected.
But I realised something:
Why should the public production application ship that entire interface at all?
A permission check protecting an action is good.
Not exposing unnecessary internal surface in the first place is better.
So production gained a stricter separation.
The private studio can be disabled entirely.
The route can behave as though it does not exist.
And the production bundle does not need to carry the internal implementation just because the development version has it.
This is a subtle change.
Nobody studying vocabulary will notice.
That is the point.
Security headers are spectacularly boring
I also spent time on something I had almost no interest in when I started this project.
HTTP headers.
A year ago, if somebody had told me:
You will eventually get excited because your study app returns the correct production security headers.
I would have been concerned for my future.
Yet here we are.
The application now has stronger production policies around what resources can load, how sensitive responses should be cached, what browsers should allow and what unnecessary framework information should be exposed.
None of that changes the colour of a button.
None of it improves my vocabulary score.
It reduces unnecessary ways the application can behave incorrectly.
That has become enough reason for me to care.
I also needed a safer way to package the project
This was another lesson from treating development and release as separate things.
My development folder contains things a public release should not.
Local configuration.
Internal tooling.
Development assumptions.
Files useful to me but irrelevant to somebody receiving a clean build.
It is surprisingly easy to think:
I'll just remember not to include those.
That is not a system.
That is a future mistake waiting for the right tired evening.
So release safety started becoming automated too.
Environment files should be excluded.
Expected templates should exist.
Build checks should fail when the archive is obviously unsafe.
The goal is to stop relying on memory for rules important enough to automate.
That principle has appeared everywhere in this project.
If forgetting one step can create a serious problem, maybe the software should remember the step instead.
The tests became less about features and more about promises
Earlier tests were satisfying because they proved features.
Can I answer a question?
Can the exam submit?
Can the app sync?
Can the vocabulary engine schedule something?
Later tests started proving less visible promises.
Does the production build succeed?
Do the existing automated checks still pass?
Did internal editor code remain outside the public production bundle?
Does the internal route disappear when it should?
Do the expected response protections exist?
Does the release package avoid carrying things it should not?
Does disabling an external integration actually disable it cleanly?
Those tests do not make the app more impressive.
They make the word working mean slightly more.
There is still a lot I am not calling finished
This part matters.
Hardening the app does not mean it is suddenly ready for everybody.
It isn't.
The current codebase is much closer to a serious private alpha or early beta than the prototype I started with.
But several things still need real-world abuse.
The local-first architecture exists.
That does not mean I have proven every ugly connectivity scenario.
I still need stronger testing around:
multiple tabs,
different devices,
account switching,
reconnections,
stale information,
interrupted writes,
storage problems,
and what happens when clocks or state disagree.
The app also does not yet have a complete cold offline experience where I can confidently launch everything from scratch without connectivity.
Native Android and Windows packages are not finished either.
And active-session recovery still deserves more real-device testing before I treat every interruption case as solved.
I would rather leave those statements here than quietly upgrade:
implemented
into:
proven.
Those words should mean different things.
“Private alpha” is not an insult
Earlier in this project, I probably would have found that label disappointing.
After building this much, I actually like it.
It says:
A lot exists.
A lot works.
The architecture has become substantial.
I can use the product seriously.
But I am still deliberately looking for reasons not to trust it yet.
That is a healthier stage than pretending a polished interface means the engineering underneath has finished.
There is a huge difference between:
I can use this
and:
I am comfortable asking somebody else to rely on this.
I am somewhere between those sentences.
The bar moved because the app moved
The first version had a wonderfully low standard.
If it loaded and gave me six questions, I was happy.
Then I wanted it to remember mistakes.
Then vocabulary.
Then progress.
Then cloud sync.
Then exam simulation.
Then personal intelligence.
Then content operations.
Then recovery.
Every capability raised the cost of being wrong.
That is the part of product development I understood least when I began.
Features do not only add value.
They add obligations.
If you store something, protect it.
If you sync something, reconcile it.
If you infer something, justify it.
If you create an account, let the owner control it.
If you build an internal tool, keep it internal.
If you depend on an external system, expect it to fail.
If you call something reliable, try very hard to prove yourself wrong first.
It worked for me
And for the beginning of this project, that was enough.
It was actually the correct standard.
There was no reason to build enterprise-grade infrastructure around 32 questions and one user.
I needed to learn whether the basic idea was useful.
It was.
Then the product grew.
The standard had to grow with it.
That does not mean I suddenly need perfection.
There is no such thing.
It means I no longer want my own familiarity with the codebase to be part of the product's reliability model.
I should not need to know which button to avoid.
Which state to manually clear.
Which error can safely be ignored.
Which file definitely should not be distributed.
Which failure only looks scary.
The software needs to carry more of that knowledge itself.
A private prototype can depend on the developer being nearby.
A real product eventually has to survive when the developer is not.
I'm not completely there yet.
But that is now the standard I'm building toward.
Because the question is no longer only:
Does this work for me?
It is becoming:
Would I trust this if I hadn't built it?
That is a much harder question.
Good.