Back to all writing

Passing Tests Was Giving Me False Confidence

IPMAT Series•Part 32•6 min read•By Mikhil

Passing Tests Was Giving Me False Confidence

Green is a dangerous colour.

Especially in a terminal.

PASS

PASS

PASS

For a while, those lines made me feel much safer than I should have.

I had tests.

Some important engine suites passed.

The Intelligence-specific tests all passed.

Quant had a lot of passing checks.

Verbal had years of accumulated validation.

Good.

Then I ran the projects as repositories instead of celebrating the nicest subset.

The confidence dropped very quickly.

A test can only prove what it tests

This is the sentence I should have respected more.

A passing test does not mean:

The product works.

It means:

This specific assertion passed under these specific conditions.

That is still useful.

It is not the same claim.

Some of my tests were pure-function tests.

Some were source checks.

Some expected filenames.

Some validated architecture.

Some did not run in the main command.

Some still referenced old migrations.

Some important real-world behaviour could not be reproduced inside the local environment at all.

The green output had flattened all those differences.

Intelligence was the clearest example

The dedicated Intelligence tests passed.

That was real.

The inference engine had useful structure.

Evidence strength.

Freshness.

Contradictions.

Planning.

Recommendations.

Retention.

Those pieces deserved credit.

But the overall repository still had failing checks and TypeScript problems.

That meant both statements were true:

The Intelligence engine has strong tested parts.

and:

The Intelligence product is not release-ready.

Software makes it very easy to accidentally replace the second sentence with the first.

Quant had the same illusion from another direction

Quant had many passing product checks.

Practice.

Review.

Analytics.

Security classification.

Syntax.

Again:

Real progress.

The production build still exposed unresolved issues.

And the much bigger product problem remained:

The architecture was far ahead of the usable question bank.

A perfect test for an empty feature does not make the feature useful.

Testing correctness and testing usefulness are different jobs.

Some tests were stale

This is almost funny.

The project evolved.

Migration names changed.

Old modules disappeared.

New architecture replaced old architecture.

Some tests kept checking the old world.

Now a failure might mean:

The product is broken.

Or:

The test is obsolete.

That is not a reason to ignore the failure.

It is a reason to repair the test suite.

A stale test still damages trust because now the signal itself is noisy.

If red does not reliably mean broken and green does not reliably mean safe, the suite stops doing its job.

Missing tests were worse than failing tests

At least a failing test tells me something.

A test that never runs can create confidence without evidence.

That happened too.

Important checks existed but were not included in the default verification command.

So I could run the normal script, see green, and still miss a known failure elsewhere.

That is a process bug.

The main validation command should represent the release gate.

Not a convenient subset.

Source-string tests can become theatre

Some automated checks were essentially verifying that certain strings or structures existed in source.

Those are useful for guarding architecture conventions.

They are not proof that the behaviour works.

A file can contain the correct function name and still behave incorrectly.

A migration can contain a policy and still not match the live database.

A route can exist and still fail on a real account.

I needed to stop treating all tests as equal evidence.

TypeScript became another independent gate

The app can look good.

Pure functions can pass.

CSS can parse.

The production build can still fail.

That is why type checking became a separate release requirement rather than something I assume another test covers.

One command does not prove everything.

The release gate needs layers.

Syntax.

Types.

Unit behaviour.

Integration.

Security.

Build.

Device.

Account isolation.

Offline.

Recovery.

Each catches a different category of mistake.

The live backend cannot be proven from a ZIP

This was another important boundary.

A source archive can show intended RLS policies.

It can show migrations.

It can show grants.

It cannot prove what is actually applied in production.

Maybe a migration was skipped.

Maybe somebody changed SQL manually.

Maybe permissions drifted.

Maybe the live database contains something the repo does not.

That means some tests require the actual hosted environment.

Source confidence and deployment confidence are different.

The same is true for Android

A TypeScript check is useful.

An emulator is useful.

A debug APK is useful.

None of them prove that a physical device will behave correctly during:

Backgrounding.

Resume.

Keyboard changes.

Network loss.

Account switching.

Storage pressure.

Accessibility services.

Real haptics.

Production signing.

The environment is part of the test.

I started treating tests as evidence instead of verdicts

This matches the rest of the project surprisingly well.

A test result is evidence.

What kind?

How broad?

How current?

What does it actually support?

What does it not support?

That is exactly how I want learner analytics to behave.

So maybe it was inevitable that the engineering process needed the same philosophy.

Do not make a stronger claim than the evidence supports.

The new release gate became much less glamorous

Before calling something complete, I now want much more than:

npm test passes.

I want:

Type checking.

All critical tests included.

Migration verification.

Security checks.

Real account isolation.

Account switching.

Offline survival.

Retry/idempotency.

Cross-device agreement.

Accessibility QA.

Production build.

Clean release archive.

Eventually payment sandbox behaviour.

Eventually real physical-device testing.

That sounds exhausting.

It is.

That is what "dependable" costs.

Tests did not fail me

This is important.

The mistake was not using tests.

The mistake was asking them to answer a larger question than they could.

A good unit test was doing exactly what it was supposed to do.

I was the one converting:

This function behaves correctly

into:

The product is safe.

That leap was mine.

Green should become boring

This is probably the goal.

I do not want a passing suite to feel like a celebration.

I want it to feel like one required condition among many.

Of course it passes.

Now what about the build?

What about staging?

What about the second account?

What about offline?

What about the phone?

What about the real database?

That mindset is less satisfying.

It is also much harder to fool.

Confidence needs coverage

The lesson sounds almost identical to the analytics article I wrote earlier.

A mastery score without enough evidence is weak.

A green test result without enough coverage is weak.

In both cases, the answer is not pessimism.

It is calibration.

Trust the evidence for what it actually proves.

Nothing more.

That has become one of the most repeated lessons in this entire project.

Apparently I needed to learn it in both education and software engineering.