Then I Closed the Tab Mid-Exam
There is an extremely convenient way to design an exam system.
Assume nothing goes wrong.
The student starts the mock.
The timer runs.
They answer questions.
They submit.
The result is saved.
Perfect.
That is also not how computers work.
Tabs get closed.
Pages refresh.
Browsers crash.
Laptops sleep.
Phones kill background applications.
Someone opens another app and comes back later.
A connection disappears.
Maybe the student simply exits without thinking about it.
And once I started treating the exam engine as something I might genuinely depend on, one question became difficult to ignore:
What happens if the session is interrupted before I finish?
The first answer was not good enough.
So I started building recovery.
Historical progress was already protected
By this point, the app had already become much more careful about preserving progress.
Practice history mattered.
Vocabulary state mattered.
Exam results mattered.
Mistakes mattered.
Bookmarks mattered.
The app had local-first behaviour, syncing, an outbox for pending work and more protection around important state.
That solved one class of problem:
What happens after something has already been recorded?
An active session is different.
It is unfinished.
The result does not exist yet.
The student may have answered only part of the exam.
The timer may still matter.
Some questions may have been visited.
Others may not.
There may be answers selected but no final submission.
That is not historical progress.
It is work in progress.
And I had not been treating those two things as separate enough.
Closing a tab should not erase forty minutes
Imagine being deep into a mock.
You have been solving for forty minutes.
You close the tab accidentally.
Or the browser crashes.
Then you reopen the app.
Nothing.
Start again.
I would hate that.
And if I would hate it, I should probably not design the product to do it.
The same problem exists in smaller practice sessions too, but an exam makes the cost much more obvious.
A three-minute practice set disappearing is annoying.
A long mock disappearing can make the entire session feel wasted.
That meant the app needed to remember more than completed attempts.
It needed some concept of an active session.
An unfinished exam is not a failed exam
This sounds obvious.
But software loves reducing messy situations into clean states.
Started.
Finished.
Passed.
Failed.
Submitted.
Not submitted.
Real use is less neat.
If I leave an exam halfway through because the browser closed, that does not necessarily mean I abandoned it.
If the app crashes, that definitely does not mean I intentionally submitted.
If I come back after an interruption, the system should not invent an outcome merely because the session stopped being visible.
So I started treating interruption as its own condition.
Not success.
Not failure.
Not submission.
Just:
This session existed and did not finish normally.
That gave recovery somewhere to live.
Practice and exams should not recover identically
This became another case where the distinction between learning and testing mattered.
I had already separated normal practice from exam simulation because they serve different purposes.
Practice is flexible.
Exams are constrained.
That difference continues even when something goes wrong.
Suppose I am casually practising vocabulary.
If I leave and come back later, restarting the session may not matter much.
The goal was learning.
There is no competitive result being protected.
A mock exam is different.
Timing matters.
Answers matter.
The sequence matters.
The conditions under which the attempt happened matter.
If recovery changes those conditions too much, the result stops meaning what I think it means.
So I did not want one generic:
Resume where you left off
feature pasted onto everything.
The recovery rules need to respect what kind of session was interrupted.
The timer became the difficult part
Answers are relatively easy to understand.
I selected option B.
Store option B.
The timer is more complicated.
Suppose I start a timed exam.
Then I close the tab.
What should happen to the clock?
Should time stop?
Should it continue?
Should the session become invalid?
Should the app ask what happened?
Should a recovered attempt still count as a normal mock?
There is no completely neutral answer.
Pausing automatically can make the exam easier than intended.
Continuing blindly can punish someone for a technical failure.
Resetting would be worse.
This is where I realised that recovery is not simply about restoring bytes.
It is about preserving the meaning of the session.
The app should not tell me I completed a strict timed mock under real conditions if the recovery behaviour secretly changed those conditions.
That would make the data look more precise than it actually is.
And I have already learned how dangerous that can be.
A recovered session needs context
This changed how I thought about persistence.
Originally, the question was:
Can I restore the session?
Eventually it became:
Can I restore it honestly?
Those are different questions.
Restoring answers is useful.
Restoring progress is useful.
But the app also needs enough context to understand what kind of interruption happened and what the recovered session now represents.
Otherwise recovery can preserve the interface while corrupting the meaning.
That would be a strange kind of reliability.
The screen survives.
The data becomes misleading.
I would rather have an honest limitation than a perfect-looking continuation that quietly changes the rules.
Normal practice created a different timing problem
Practice sessions also record timing.
That is useful.
Response time can reveal hesitation.
It can help identify questions that are technically correct but painfully slow.
It can show whether speed is improving.
But a timer running while the tab is abandoned tells me almost nothing about learning.
If I answer a question, walk away for twenty minutes and return, the app should not confidently interpret that as:
Response time: 20 minutes.
Technically, that might be what the clock observed.
Educationally, it is garbage.
That reminded me of a rule I keep encountering:
Recorded data is not automatically meaningful data.
Software can measure something perfectly and still draw the wrong conclusion from it.
So interruption recovery also affects analytics.
Not only UX.
I needed active state to become explicit
The earliest versions of the app could get away with being simple.
Open a set.
Answer things.
Save the result.
Move on.
As the product became more sophisticated, implicit state became dangerous.
If the system needs to recover something later, it first needs to understand what actually existed.
Which session?
Which mode?
Which questions?
Which answers?
Was it finished?
When did it begin?
What state was it in when the interruption happened?
I am deliberately not turning this article into a description of the internal implementation.
That part will keep evolving anyway.
The important shift was conceptual.
An active learning session became something worth modelling deliberately rather than something that only existed because a page happened to be open.
Recovery also needed to avoid creating duplicates
There was another boring problem hiding underneath this.
Suppose the student returns.
The app restores something.
Then synchronization happens.
Then another piece of code sees the same session.
Now what?
I did not want recovery to create two histories for one attempt.
Or two submissions.
Or one local attempt and one cloud attempt that later look like different events.
Or a restored exam that silently fights with an already-finished version somewhere else.
This is the kind of problem that never appears in a product screenshot.
There is no exciting animation for:
We successfully did not create duplicate state.
But avoiding that matters.
A recovery system that occasionally invents extra attempts would be worse than having no recovery at all.
So the feature had to fit into the reliability work I had already been doing rather than exist as a completely separate patch.
The app should not punish someone for using it normally
This became the simplest rule.
Refreshing a webpage is normal.
Closing a tab is normal.
Switching applications is normal.
Devices sleeping is normal.
Internet connections disappearing is normal.
Software failing occasionally is also normal.
The user should not need to behave like a systems administrator to protect their studying.
I do not want someone thinking:
I better not touch anything or I might lose the session.
The app should carry more of that responsibility.
Obviously there are limits.
No recovery system can guarantee that nothing will ever be lost.
Devices can fail completely.
Storage can be cleared.
Browsers can behave unexpectedly.
Bugs exist.
But there is a huge difference between acknowledging those limits and designing as though interruption never happens.
This feature is almost invisible when it works
That is becoming a recurring theme in this project.
The features that make the app look sophisticated are easy to show.
Vocabulary modes.
Exam dashboards.
Heatmaps.
Recommendations.
Progress graphs.
AI feedback.
The features that make the app trustworthy often look like nothing.
A pending change waits until the internet returns.
A duplicate is not created.
A session remembers enough to recover.
A stale state does not silently overwrite something newer.
An interrupted timer is not treated as clean learning evidence.
Nothing dramatic happens.
That is the success.
I keep discovering the difference between a demo and a product
A demo assumes cooperation.
The browser stays open.
The network works.
The user follows the intended path.
Every request finishes.
Nobody refreshes at the wrong moment.
Nobody leaves halfway through.
Under those conditions, almost everything feels finished.
A product has to survive people behaving normally.
That includes doing things the developer did not imagine when designing the happy path.
I had already learned this while building offline sync.
I had already learned it while building SYNTHESIS.
Now the exam engine was teaching me the same lesson again.
A successful path proves that something can work.
The annoying paths determine whether I trust it.
I tested the wrong thing at first
When I originally built exam mode, most of my attention went toward whether the exam itself behaved properly.
Did the right questions appear?
Did the timer work?
Did marking work?
Did submission work?
Did the result look right?
Those were reasonable tests.
But they all shared one assumption:
I completed the session normally.
Once that worked, the more interesting tests became destructive.
What happens if I leave?
What happens if I refresh?
What happens if the session is unfinished?
What happens if some state exists locally but not elsewhere?
What happens if I come back later?
What happens if the system has to decide whether something should still be resumed?
Breaking the expected path revealed more about reliability than repeating the successful path ever could.
Recovery should not become permission to lie
There is a tempting way to make recovery feel amazing.
Restore everything.
Hide every interruption.
Make it look as if nothing happened.
For ordinary practice, that can sometimes be exactly what I want.
For strict exam simulation, I am more careful.
If something materially changed the conditions of the attempt, the app should not quietly pretend otherwise.
A mock result is useful partly because I trust the conditions that produced it.
If recovery makes that ambiguous, the product should preserve that ambiguity rather than manufacture certainty.
This is the same principle behind the personal-intelligence system.
Two wrong answers should not confidently diagnose a weakness.
A small amount of evidence should remain a small amount of evidence.
And an interrupted exam should not become a perfectly ordinary exam merely because the UI recovered successfully.
Reliability keeps moving upward
At first, reliability meant:
Don't lose my answers.
Then:
Don't lose my progress.
Then:
Don't lose changes when I'm offline.
Then:
Don't let sync create corrupted state.
Now:
Don't lose the thing I'm currently doing.
Every time the app becomes more capable, another layer becomes worth protecting.
That is probably unavoidable.
The first 32-question version did not need sophisticated recovery because there was very little at stake.
The current app remembers enough about me that losing or misinterpreting state can affect everything built on top of it.
That makes reliability cumulative.
The smarter the system becomes, the more carefully it has to remember what actually happened.
Then I closed the tab
And for once, closing the tab became part of the test.
Not an accident outside the test.
The app should expect interruptions.
It should expect imperfect networks.
It should expect devices to disappear.
It should expect students to leave.
It should expect the real world to interfere with the clean sequence I designed on my laptop.
That does not make the product less elegant.
I think it makes it more honest.
A timer that works for sixty uninterrupted minutes is easy.
A learning system that can survive interruption without inventing evidence, duplicating progress or pretending the conditions never changed is much more interesting.
The exam engine had already learned how to start a mock.
Now I wanted it to understand that sometimes the student disappears before the mock ends.
And when they come back, the app should have a better response than:
What exam?