The first working version of SYNTHESIS did not look like JARVIS.
It had no voice.
It could not see my screen, remember our conversations, manage my applications, or independently complete a complicated task.
It lived inside a terminal and printed a few lines of text.
But for the first time, the project did something that mattered.
It registered two tools. It recognised that one was harmless and the other was dangerous. It executed the harmless one automatically, stopped before the dangerous one, asked for permission, and respected my refusal.
That small test was the first real proof that SYNTHESIS could become a system rather than remain an idea.
Starting with empty folders
The project began with a basic Python environment on Windows and a collection of folders representing the systems I expected SYNTHESIS to need:
agenttoolsmemorymodelsperceptioninterfacetestsconfig
There were also spaces for the application itself, shared core logic, a main entry point, requirements, environment configuration, and documentation.
Most of those folders were empty.
Seeing them together created the appearance of a large architecture, but empty folders do not create an assistant.
I needed one complete path through the system.
Something small enough to understand and test, but important enough to prove that the architecture could work.
That became the first vertical slice.
What a vertical slice meant
A vertical slice is a narrow piece of a system that works from beginning to end.
Instead of partially building perception, memory, voice, tools, and an interface simultaneously, I wanted one request to travel through the essential layers successfully.
For SYNTHESIS, that path looked roughly like this:
- A model produces a structured response.
- The system identifies a requested tool.
- It finds that tool in a registry.
- A permission manager evaluates the risk.
- An executor runs the tool only when permitted.
- The result is verified.
- The system reports what happened.
This was not yet an autonomous agent.
It was the smallest useful skeleton of one.
Beginning without a real AI model
SYNTHESIS eventually needs intelligence capable of interpreting natural-language requests.
But connecting a real model immediately would have made the first tests harder to understand.
If something failed, I would not know whether the problem came from the model, the prompt, the response format, the tool, the executor, or the permission system.
So the first version used a mock model adapter.
The purpose of the adapter was to create a standard boundary between SYNTHESIS and whichever model might eventually provide its reasoning.
The mock version returned predictable responses. It did not need an API, internet connection, or expensive hardware.
That predictability made it useful.
A mock system is not impressive, but it removes uncertainty. It allowed me to test the architecture without pretending the intelligence layer was already complete.
It also meant the project was not permanently tied to one model provider from its first day.
Defining what a tool is
A personal agent becomes useful when it can act through tools.
But SYNTHESIS could not treat every Python function as an uncontrolled ability. It needed a common definition of what a tool was.
Each tool required a name, a description, an assigned risk level, an execution method, a verification method, and a structured description of the information it accepted.
That shared structure gave the rest of the system a predictable way to interact with different abilities.
The first tool was deliberately harmless: retrieve the current time.
The second was deliberately labelled as dangerous and existed only to test the permission boundary.
The purpose was not the actions themselves.
It was whether SYNTHESIS could treat them differently.
The registry answered one question
As tools are added, the system needs to know which ones exist.
That became the responsibility of the tool registry.
A tool could be registered, retrieved by name, checked for existence, listed, and represented in a standard format. Duplicate names were rejected because the system should never be uncertain about which ability it is invoking.
This was simple infrastructure.
But it established an important boundary: the model could not invent an ability and silently execute it.
SYNTHESIS could only use tools that the surrounding system had deliberately registered.
The model might request an action.
The application still controlled whether that action existed.
Giving actions different levels of risk
Not every action deserves the same amount of friction.
Retrieving the time is harmless. Constantly asking for confirmation would make the assistant frustrating without making it meaningfully safer.
Other actions could affect files, accounts, messages, money, privacy, or the state of the computer. Treating those actions like reading the clock would be irresponsible.
The first permission system therefore classified tools into three levels:
- Low risk
- Medium risk
- High risk
Low-risk actions could execute automatically.
Medium-risk actions could follow a configurable policy.
High-risk actions required explicit user confirmation.
This was still a basic model. Real-world risk depends on more than a label attached to a tool.
But the first test needed to prove one principle:
The model requesting an action does not mean the action is automatically authorised.
Execution was not enough
The tool executor connected the registry and the permission manager.
It located the requested tool, checked the permission decision, requested confirmation when necessary, performed the action, and returned a structured result.
But calling a tool successfully did not prove that the intended outcome occurred.
That is why tools also needed verification.
An application can accept a command without completing it. A file operation can fail. A request can return an unexpected result. A tool can run without producing the outcome the assistant claims.
SYNTHESIS should distinguish between:
I attempted the action
and:
I verified that the action succeeded.
The first vertical slice could only perform simple verification, but establishing the requirement early mattered.
An assistant should not manufacture confidence after execution.
The first test failed
I initially ran the test file directly:
python tests/test_tools.py
It failed with a ModuleNotFoundError.
Python could not find the project’s tools module because of how the test had been launched.
The test needed to be executed as a module from the project root:
python -m tests.test_tools
After making that correction—and saving one file I had accidentally left unsaved—the test ran.
That small failure was useful because it made the project feel real.
SYNTHESIS was no longer only a plan being discussed. It was a codebase capable of failing for specific reasons that could be investigated and corrected.
The first successful run
The terminal displayed two registered tools.
The low-risk time tool was allowed automatically. It executed successfully, returned the current time, passed verification, and reported success.
Then the test reached the dangerous tool.
The permission manager did not run it automatically. It displayed a security gate and asked for confirmation.
I entered n.
The executor refused the action and returned:
Execution denied by user.
Nothing dramatic happened.
That was the success.
The system demonstrated that it could possess an ability without automatically exercising it.
It also demonstrated that a denial was not an error to bypass. It was a valid outcome that the assistant needed to respect.
What this tiny test proved
The first vertical slice did not prove that SYNTHESIS was intelligent.
It did not prove that the architecture would scale, that future models would make reliable decisions, or that a three-level permission system could handle every real situation.
It proved something narrower:
The basic pieces could communicate.
A model response could be normalised. Tools could be registered. Risk could influence execution. Dangerous actions could stop for confirmation. Results could be verified and reported.
That foundation was more valuable than beginning with a voice and pretending the hard parts had already been solved.
A cinematic interface might have made SYNTHESIS look alive.
The vertical slice made it structurally real.
Small systems reveal large questions
Even this tiny test exposed questions that later versions must answer.
Who assigns a tool’s risk level?
Can a harmless tool become dangerous because of its arguments?
What happens when several low-risk actions combine into a high-risk outcome?
How should the assistant explain what it wants to do?
How can verification remain independent from the action being verified?
What should happen when the user refuses one step inside a larger plan?
The first architecture did not answer all of them.
It gave those questions somewhere concrete to live.
SYNTHESIS had moved from:
What if I built a personal AI assistant?
to:
Here is one action travelling through a system that can decide whether it is allowed.
That is a much smaller sentence.
It is also the point where building truly began.