I Didn’t Want Another Dashboard Full of Numbers
There is something extremely satisfying about an analytics dashboard.
Big percentage.
Green arrow.
Pretty graph.
Heatmap.
Mastery score.
A line moving upward.
Put enough of those on one page and a product immediately looks intelligent.
That is also what made me suspicious of them.
Because the easiest thing in the world would have been to make Quant look like it understood me long before it actually did.
The deeper I got into Quant, the more I realised that collecting numbers was not the difficult part.
The difficult part was deciding what I was allowed to conclude from them.
Accuracy is useful
If I answer 100 questions and get 80 correct, that is useful.
But then the questions start.
Which questions?
Were they all easy?
Were some repeats?
Were they spread across the syllabus?
Timed or untimed?
Did I answer correctly on the first try?
Did I change my answer?
Was I confident?
Suddenly:
80%
contains much less information than it appeared to.
That does not make accuracy useless.
It means accuracy is an observation.
Not a complete diagnosis.
Time has the same problem
Ninety seconds is ninety seconds.
Except the meaning changes with the question.
45 seconds on an easy percentage problem is not the same as 45 seconds on a hard algebra problem.
20 seconds and wrong may mean rushing.
Six minutes and correct may mean conceptual understanding but poor exam execution.
So one of the earliest rules became:
Never interpret speed by itself.
Compare like with like
Suppose the dashboard says:
You're 18 seconds faster than last week.
Nice.
Except last week I solved hard Geometry.
This week I solved easy Arithmetic.
Am I faster?
Maybe.
The data has not proved it.
So comparisons need to control for things like topic, difficulty, format, exam context and timed versus untimed conditions.
Not perfectly.
This is not a laboratory.
But enough that the comparison means something.
Median time became more useful than average time
One outlier can destroy an average.
Maybe I got distracted.
Maybe one question genuinely took forever.
Maybe I left the tab.
A median often describes normal pace better.
The goal is not more statistics.
It is statistics that distort reality less.
Speed and accuracy tell a better story together
A topic can fall into four rough states.
Fast + Accurate — probably strong.
Fast + Inaccurate — control may be the problem.
Slow + Accurate — understanding may be fine, execution needs work.
Slow + Inaccurate — deeper weakness may exist.
That already tells me more than:
Algebra accuracy: 68%
because 68% can hide completely different problems.
Difficulty complicates everything again
Suppose one learner has 90% accuracy on mostly easy questions.
Another has 74% on harder material.
Which has stronger mastery?
Accuracy alone cannot answer.
But difficulty itself is initially an editorial judgment.
A label is useful.
It is not divine truth.
Real learner evidence may later calibrate it.
Until then, difficulty needs uncertainty too.
Mastery became a dangerous word
If the app says:
Quadratics Mastery: 76%
what does that mean?
There is no mastery sensor inside the learner's brain.
It is an estimate from evidence.
So I did not want mastery to be correct answers divided by total answers with a nice ring around it.
The model can consider accuracy, difficulty, pace, recency, breadth, retention and confidence.
The most important part is that the conclusion can be explained.
Why does the app think this topic is strong?
Why does it think another is weak?
If the answer is only:
Algorithm.
that is not good enough.
Evidence strength matters
Imagine:
Arithmetic mastery: 82%, evidence strong.
Geometry mastery: 91%, evidence low.
Which conclusion should I trust more?
Probably Arithmetic.
That is why the analytics architecture includes evidence bands.
None.
Low.
Moderate.
Strong.
The higher number is not automatically the stronger conclusion.
Sometimes the correct result is nothing
A new learner opens Analytics.
There is not enough evidence.
What should the page show?
Software wants to fill space.
Zero scores.
Sample charts.
Fake insights.
I would rather show:
Not enough evidence yet.
If there are not enough comparable periods, do not claim improvement.
If there is not enough timed evidence, do not claim readiness.
If a reference threshold has not been verified, do not invent one.
Empty is sometimes the most intelligent state.
Readiness is not mastery
Mastery asks:
Can I solve this?
Readiness asks:
Can I execute under the conditions that matter?
A learner can understand Quadratics and still be too slow under exam pressure.
The concept may not be the problem.
Execution may be.
So readiness uses different evidence.
And it should not become a fake score prediction.
I do not want:
You are 81% ready.
unless that number has an extremely clear meaning.
A confident-looking percentage can affect how somebody studies.
The burden of honesty should be higher.
Mock analysis reveals different behaviour
A mock tells me things ordinary practice cannot.
Which questions did I attempt?
Which did I skip?
Where did I lose time?
Where did negative marks come from?
Which easy opportunities did I miss?
Where did I overinvest?
The system can classify patterns like efficient attempts, overinvestment, good skips and missed opportunities.
But those labels are still inferences.
The application cannot read my mind.
The confidence of the sentence should match the confidence of the evidence.
Then we started collecting evidence before the final answer
At first an attempt could tell me:
Question.
Answer.
Correctness.
Time.
Then the evidence model got much richer.
When was the question shown?
When did I first interact?
What was my first answer?
Did I change it?
How many times?
Did I clear a Short Answer response?
Did I leave and return?
How much time was active?
How much was idle?
How confident was I before seeing the result?
Did I open the explanation?
Had I seen the question before?
Where in the session did it appear?
For DI, how much time was spent interpreting the set before solving?
These signals help distinguish behaviours that final accuracy collapses together.
First answer versus final answer is fascinating
Suppose I choose A.
Then change to C.
Then submit C.
C is correct.
Normal analytics says:
Correct.
Done.
But maybe answer changes usually help me.
Maybe they usually hurt me.
One changed answer proves almost nothing.
A repeated pattern may.
Observation first.
Interpretation later.
Confidence needs to be captured before feedback
If I ask:
How confident were you?
after showing:
Correct!
the answer is contaminated.
So pre-reveal confidence creates a better signal.
High confidence + wrong may indicate something different from low confidence + wrong.
Across enough attempts, the app can start learning whether the learner is calibrated about their own knowledge.
That is a very different skill from raw accuracy.
Active time needed to separate from elapsed time
If a question sits on screen for six minutes, that does not prove six minutes of solving.
Maybe the learner switched tabs.
Maybe they got distracted.
The app does not need to know what they did elsewhere.
It only needs enough restraint not to call all six minutes problem-solving time.
Better intelligence sometimes comes from making fewer assumptions.
Previous exposure changes the meaning of correctness
If I answer the same question correctly four times, that is not equivalent to four independent correct questions.
By the fourth attempt, I may remember the answer.
That can be useful evidence for retention.
It is weak evidence for general topic mastery.
So exposure and retry interval matter.
I solved a new question
and
I remembered an old answer
should not be treated identically.
Explanation views are not mastery
Opening an explanation is an event.
Not proof of repair.
The useful question is what happens later.
Does the learner solve another related problem?
Does the same mistake return?
The repair is not:
Explanation opened.
The repair is:
Future evidence suggests the learner can now do the thing.
Visuals still matter
After all of this complaining about dashboards, I added a lot of visuals.
Mastery radar.
Evidence maps.
Performance trends.
Speed-versus-accuracy plots.
Mistake distributions.
Activity heatmaps.
I like them.
The rule is simple:
The picture does not get to become more confident than the data underneath it.
A beautiful chart built from three questions is still a bad chart.
Even a time-range filter can lie
Suppose Analytics offers 7 days, 30 days, 90 days and all time.
If I select 7 days, I naturally assume the page now describes seven days.
But if one chart uses the selected range while another calculation quietly uses all historical progress, the page becomes internally inconsistent.
Every number may be mathematically correct.
The dashboard is still misleading.
That is the kind of issue the broader audit exposed and it still needs cleanup.
More analytics can make the product worse
At one point Home contained almost everything.
Focus.
Momentum.
Metrics.
Trends.
Mastery.
Evidence.
Achievements.
Review.
Learning summaries.
Useful individually.
Exhausting together.
The learner usually has a simpler question:
What should I do now?
So deeper analysis should exist when I want it.
It should not be forced into every study session.
A recommendation is also an analytical claim
If the app says:
Practice Percentages next
it is claiming that this is probably a useful use of my time.
That recommendation should have evidence.
Due review.
Unrepaired mistakes.
Weakest measurable topic.
Or maybe not enough history yet.
The learner should be able to understand why.
The real question became: what does the app know?
A wrong answer happened.
Knowable.
A response took 74 active seconds.
Knowable.
The learner changed from B to D.
Knowable.
The learner has a conceptual weakness.
Inference.
The learner struggles under pressure.
Inference.
The learner is exam-ready.
Much stronger inference.
The further the system moves from observation, the more evidence I want underneath it.
That hierarchy now shapes almost every intelligent feature.
I still like the graphs
I am not against dashboards.
I am against dashboards pretending to know more than they do.
The visuals make the product feel alive.
When the evidence underneath them is real, they are genuinely useful.
The goal is not maximum measurement.
It is useful measurement with honest limits.
So what should analytics actually do?
It should help me understand:
What is happening?
What is changing?
How confident are we?
Why might it be happening?
What should I do next?
And sometimes:
We don't know yet.
That last sentence may be the most important analytical feature I build.
The charts are just the interface.
The hard part is earning the sentence underneath them.